Embedding Apache Spark in an Application
June 18, 2018
Recently I wanted to do some data encryption on a data set while preserving the format. After looking at few things I settled on this library which implements FF1 encryption standard.
Since this is a Java library I chose to use Scala. The data was a CSV dump from Postgres with headings and Spark has great support for reading and writing large CSV data sets. Furthermore I could use all the cores on the machine to speed up processing without writing code to parallelize the encryption.
One of the issues with deploying a spark application is setting up the Spark environment. The dataset itself wasn’t big enough for a cluster (about 40 million rows).
Since Spark itself is a bunch of jar files I was thinking of a simpler solution. I was using Gradle as the build tool (it has good Scala support) so it is trivial to package the Spark runtime and Scala runtime as dependencies using Gradle application plugin.
dependencies {
compileOnly("org.scala-lang:scala-library:2.11.8")
compileOnly("org.scala-lang:scala-reflect:2.11.8")
compileOnly("org.scala-lang:scala-compiler:2.11.8")
compile 'org.apache.spark:spark-sql_2.11:2.1.0'
compile 'org.apache.spark:spark-launcher_2.11:2.1.0'
compile 'org.apache.spark:spark-catalyst_2.11:2.1.0'
compile 'org.apache.spark:spark-streaming_2.11:2.1.0'
compile 'org.apache.spark:spark-core_2.11:2.1.0'
compile 'com.idealista:format-preserving-encryption:1.0.0'
}From then on it was not much different to developing a typical Spark job. Using Gradles’ JavaExec support it is trivial to locally test the application.
Written by Francois Fernando, a software craftsman and tinkerer.