Does data locality matter any more?
September 20, 2017
Yesterday went to the Spark Meetup in Sydney. One of the speakers Mike Seddon mentioned an important point. Most distributed processing systems including Hadoop emphasize data locality - the advantage of performing the computation where the data resides.
I have also started reading about Greenplum as that is what IAG is using for the the data warehouse. Greenplum also mentions about moving the computation to the data.
But most of the complex use cases involve joining multiple tables and it would be hard to imagine that the tables have been distributed in such a way that all the data partitions required for joining are co-located.
In fact one of the colleagues who have worked on Greenplum mentioned that not distributing tables according to the join patterns is a performance problem in Greenplum.
So I think for now I have agree with Mike Seddon that data locality doesn’t matter in practice.
Written by Francois Fernando, a software craftsman and tinkerer.