
Apache Flink
Hadoop
Apache Kafka
Apache Hive
Apache Storm
Splunk
Apache Airflow
Apache Spark is an engine for big data processing, with built-in modules for streaming, SQL, machine learning and graph processing.

Amazon EMR
Google BigQuery
Qubole
Snowflake
Databricks
Apache Beam
Amazon Kinesis
Google Cloud Dataflow is a fully-managed cloud service and programming model for batch and streaming big data processing.

Which is more popular?
Based on our record, Apache Spark should be more popular than Google Cloud Dataflow. It has been mentioned 80 times since March 2021.
Website, pricing, platforms and company facts side by side.
|
|
|
|
|---|---|---|
| Website | spark.apache.org | cloud.google.com |
| Pricing | — | |
| Listed in |
What each product offers, as listed by its team.


Possible disadvantages
Possible disadvantages
An editorial look at what each product does well and who it suits.


Overall verdict
Why this product is good
Recommended for
Overall verdict
Why this product is good
Recommended for
Walkthroughs and reviews on video.
Weekly Apache Spark live Code Review -- look at StringIndexer multi-col (Scala) & Python testing
More videos
Introduction to Google Cloud Dataflow - Course Introduction
More videos
How often each product is chosen within a category, 0–100% relative to the other.


Share your experience with using Apache Spark and Google Cloud Dataflow. For example, how are they different and which one is better?
External articles and on-site reviews we used to compare the two products.


Apache Spark is an open source data processing and analytics engine that can handle large amounts of data -- upward of several petabytes, according to proponents. Spark's ability to rapidly process data has fueled...
Apache Spark is a well-known, general-purpose, open-source analytics engine for large-scale, core data processing. It is known for its high-performance quality for data processing – batch and streaming with the help...
Apache Spark is an open-source and flexible in-memory framework which serves as an alternative to map-reduce for handling batch, real-time analytics and data processing workloads. It provides native bindings for the...
Google Cloud Dataflow is highly focused on real-time streaming data and batch data processing from web resources, IoT devices, etc. Data gets cleansed and filtered as Dataflow implements Apache Beam to simplify...
Recommendations tracked on public social media and blogs since March 2021.


Feature transformations should be deterministic: The same input should produce the same output when the same feature definition and configuration are applied. This is what allows training, backtesting, and live inference to remain... - Source: dev.to / 4 months ago
Apache Spark provides distributed in-memory data processing and is the appropriate tool when the data set to be reconciled does not fit in a single machine's memory, or when parallelizing the comparison across a cluster would reduce... - Source: dev.to / 4 months ago
When IoTDB was initiated in 2011, almost all influential distributed systems and databases were built in Java or on the JVM—such as Hadoop, HBase, Spark (Scala on JVM), Cassandra, Kafka, and Flink. To integrate deeply with the big data... - Source: dev.to / 6 months ago
Imo if you are using the cloud and not doing anything particularly fancy the native tooling is good enough. For AWS that is DMS (for RDBMS) and Kinesis/Lamba (for streams). Google has Data Fusion and Dataflow . Azure hasData Factory if... Source: over 3 years ago
This sub is for Apache Beam and Google Cloud Dataflow as the sidebar suggests. Source: almost 4 years ago
I am pretty sure they are using pub/sub with probably a Dataflow pipeline to process all that data. Source: almost 4 years ago
When comparing Apache Spark and Google Cloud Dataflow, you can also consider the following products.

Flink is a streaming dataflow engine that provides data distribution, communication, and fault tolerance for distributed computations.
Compare Apache Flink to Apache Spark or Google Cloud Dataflow:

Amazon Elastic MapReduce is a web service that makes it easy to quickly process vast amounts of data.
Compare Amazon EMR to Apache Spark or Google Cloud Dataflow:

Open-source software for reliable, scalable, distributed computing
Compare Hadoop to Apache Spark or Google Cloud Dataflow:

A fully managed data warehouse for large-scale data analytics.
Compare Google BigQuery to Apache Spark or Google Cloud Dataflow:

Apache Kafka is an open-source message broker project developed by the Apache Software Foundation written in Scala.
Compare Apache Kafka to Apache Spark or Google Cloud Dataflow:

Qubole delivers a self-service platform for big aata analytics built on Amazon, Microsoft and Google Clouds.
Compare Qubole to Apache Spark or Google Cloud Dataflow: