Apache Spark vs. TensorFlow

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Spark
Score 8.9 out of 10
N/A
Apache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.N/A
TensorFlow
Score 7.7 out of 10
N/A
TensorFlow is an open-source machine learning software library for numerical computation using data flow graphs. It was originally developed by Google.N/A
Pricing
Apache SparkTensorFlow
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Apache SparkTensorFlow
Free Trial
NoNo
Free/Freemium Version
NoNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache SparkTensorFlow
Considered Both Products
Apache Spark
Chose Apache Spark
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only.
Chose Apache Spark
Apache Spark is a fast-processing in-memory computing framework. It is 10 times faster than Apache Hadoop. Earlier we were using Apache Hadoop for processing data on the disk but now we are shifted to Apache Spark because of its in-memory computation capability. Also in SAP …
Chose Apache Spark
Other teams used to work on Apache Hadoop but our team started with Apache Spark directly.
Chose Apache Spark
There are a few alternatives that can do the same transformation and aggregation like Apache Spark can do but most of them are not able to perform parallel computation. For example, pandas is a really good tool to do that but not parallelized; However, there are some tools that …
Chose Apache Spark
  • Apache Spark works in distributed mode using cluster
  • Informatica and Datastage cannot scale horizontally
  • We can write custom code in spark, whereas in Datastage and Informatica we can only choose the different features proivided already.
Chose Apache Spark
Apache Spark has much more better performance and features if we compare with Hive or map/reduce kind of solutions. Spark has many other features for machine learning, streaming.
Chose Apache Spark
Spark is simply awesome to work on with any data sets and also has an in-memory database which makes it very flexible.
Chose Apache Spark
1. Apache Spark is almost 100 % faster than Hadoop.
2. Apache Spark is more stable than Amazon EMR.
3. The end to end distributed machine library is more robust in Apache Spark.
Chose Apache Spark
Databricks uses Spark as a foundation, and is also a great platform. It does bring several add-ons, which we did not feel needed by the time we evaluated - and haven't needed since then. One interesting plus in our opinion was the engineering support, which is great depending …
Chose Apache Spark
It is easy to learn, read and to maintain. It brings the best of the Ruby on Rails framework from Java that helps to create a web service so easily. Communication is one of the most distinctive features of Apache Spark compared to alternative products. You are able to …
Chose Apache Spark
We evaluated SAS alongside with Apache Spark but during the course of proof of concept found that Apache Spark was able to support the hadoop eco-system and hadoop file system much better. It was much faster at that time while having the ability to process data quickly for the …
Chose Apache Spark
I prefer Apache Spark compared to Hadoop, since in my experience Spark has more usability and comes equipped with simple APIs for Scala, Python, Java and Spark SQL, as well as provides feedback in REPL format on the commands. At the same time, Apache Spark seems to have the …
Chose Apache Spark
All the above systems work quite well on big data transformations whereas Spark really shines with its bigger API support and its ability to read from and write to multiple data sources. Using Spark one can easily switch between declarative versus imperative versus functional …
Chose Apache Spark
Even with Python, MapReduce is lengthy coding. Combination of Python with Apache Spark will not only shorten the code, but it will effectively increase the speed of algorithms. Occasionally, I use MapReduce, but Apache Spark will replace MapReduce very soon. It has many …
Chose Apache Spark
vs MapRedce, it was faster and easier to manage. Especially for Machine Learning, where MapReduce is lacking. Also Apache Storm was slower and didn't scale as much as Spark does. Spark elasticity was easier to apply compared to storm and MapReduce.
managing resources for …
Chose Apache Spark
We specifically choose Spark over MapReduce to make the cluster processing faster
Chose Apache Spark
Spark in comparison to similar technologies ends up being a one stop shop. You can achieve so much with this one framework instead of having to stitch and weave multiple technologies from the Hadoop stack, all while getting incredibility performance, minimal boilerplate, and …
Chose Apache Spark
Apache Pig and Apache Hive provide most of the things spark provide but apache spark has more features like actions and transformations which are easy to code. Spark uses optimization technique as we can select driver program and manipulate DAG (Directed Acyclic Graph)
Python …
Chose Apache Spark
There are a few newer frameworks for general processing like Flink, Beam, frameworks for streaming like Samza and Storm, and traditional Map-Reduce. I think Spark is at a sweet spot where its clearly better than Map-Reduce for many workflows yet has gotten a good amount of …
Chose Apache Spark
Spark has primarily replaced my use of writing pure Hadoop MapReduce or Apache Pig jobs for processing data. I like the fact that I can alternate between the main programming languages that I know - Java and Python - and use those to learn the Scala API. Spark also can be …
TensorFlow
Chose TensorFlow
I prefer Pytorch overall, recent models are often only available with pytorch
PyTorch is also easier to use and it is often easier to find support for PyTorch code nowadays than TensorFlow
Also it seems like lots of Google internal resource uses Jax. I mostly uses TensorFlow to …
Chose TensorFlow
TensorFlow has better support for Java compared to PyTorch and is also very well documented.
Chose TensorFlow
Can't seem to choose any deep learning platform in the above, so I'll list it here:
1. Apache MXNet: this has been used for one of our main algorithms for search as an end-to-end pipeline. We chose this because of the Scala bindings, which makes it easier to integrate with out …
Chose TensorFlow
TensorFlow provides a wide range of algorithms with more detail and customization options compared to others. Also, the library is advanced and updates regularly for optimization and new functions.
Chose TensorFlow
Most of the machine learning platforms these days support integration with R and Python libraries. So, the use of reusable libraries is not an issue. TensorFlow performs well in cloud hosting and support for GPU/TPU. However, where it lacks compared to Azure is a graphical …
Chose TensorFlow
Thought about alternatives like scikit-learn, xgboost, pytorch, caffe2, fastai exist, but they don't offer as many tools and functionality as TensorFlow does. It is better to inanest in a eco-system which is very active and well maintained by giants. Being open source, one can …
Chose TensorFlow
Keras is built on top of TensorFlow, but it is much simpler to use and more Python style friendly, so if you don't want to focus on too many details or control and not focus on some advanced features, Keras is one of the best options, but as far as if you want to dig into more, …
Chose TensorFlow

Theano is a Python library and is good for making algorithms from scratch. It is an alternative to Tensor flow. We used tensor flow because it is open source Java source and easy to learn and use.

TensorFlow is developed and maintained by Google. It's the engine behind a lot of …

Chose TensorFlow
There are lots of competitors with this library, but I think TensorFlow is the best thing for deep learning. Although it has a sharp learning curve, it's worth learning. It easy to deploy its model on Android. Keras is very good option too it, easy. In Keras, writing the neural …
Chose TensorFlow
I have used keras and matlab along with this. Also used Caffe and pyTorch sometimes, but all of them are not as powerful as TensorFlow. Keras is in good competition with TensorFlow but Keras won't allow you a lot of customization in your algorithms. And TensorFlow gives you the …
Chose TensorFlow
One major advantage of TensorFlow over Keras and other deep learning libraries is that it is the most powerful. It gives you power to write your own full customised algorithm that is not available in Keras. And it is fast too as compared to another tool as it can perform better …
Chose TensorFlow
I have used Theano to develop machine learning models, like writing the neural network. TensorFlow has reinforcement learning support and lot more algorithms while Theano does come with lots of prebuilt tools. TensorFlow provides data visualisation tools and it is possible to …
Best Alternatives
Apache SparkTensorFlow
Small Businesses

No answers on this topic

InterSystems IRIS
InterSystems IRIS
Score 8.1 out of 10
Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
Posit
Posit
Score 10.0 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 8.6 out of 10
Posit
Posit
Score 10.0 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache SparkTensorFlow
Likelihood to Recommend
9.0
(0 ratings)
6.0
(0 ratings)
Likelihood to Renew
10.0
(0 ratings)
-
(0 ratings)
Usability
8.0
(0 ratings)
9.0
(0 ratings)
Support Rating
8.7
(0 ratings)
9.1
(0 ratings)
Implementation Rating
-
(0 ratings)
8.0
(0 ratings)
User Testimonials
Apache SparkTensorFlow
Likelihood to Recommend
Apache Spark has rich APIs for regular data transformations or for ML workloads or for graph workloads, whereas other systems may not such a wide range of support. Choose it when you need to perform data transformations for big data as offline jobs, whereas use MongoDB-like distributed database systems for more realtime queries.
Read full review
  1. Whenever the problem has the demand for a neural networks based solution, Tensorflow (TF) is a great fit.
  2. The tf.dataset API makes it really simple to create complex data pipelines in a few lines of code.
  3. tf.estimators API abstracts all the complex computation graph creation logic making it very simple to get started.
  4. Eager execution makes it simple to develop a TF graph as debugging the code would be like any other imperative Python program.
  5. TF abstracts all the complexities of scaling it to multiple machines. It has various code and data distribution algorithms ready to use.
  6. Projects like TensorBoard make monitoring the training process really easy. It also gives the ability to view embeddings without any extra code. Their What-If is extremely useful for poking and understanding a black box model. It also has tools to visualize data to quickly check for anomalies.
  7. TF Autograph aims to covert any normal Python code into a distributed program which is quite handy to scale an existing code base.
Read full review
Pros
  • It performs a conventional disk-based process when the data sets are too large to fit into memory, which is very useful because, regardless of the size of the data, it is always possible to store them.
  • It has great speed and ability to join multiple types of databases and run different types of analysis applications. This functionality is super useful as it reduces work times
  • Apache Spark uses the data storage model of Hadoop and can be integrated with other big data frameworks such as HBase, MongoDB, and Cassandra. This is very useful because it is compatible with multiple frameworks that the company has, and thus allows us to unify all the processes.
Read full review
  • Data pipeline implementation is quite good, loading large amounts of data and pre-process it in an efficient way is no more issue for us
  • It supports all major DL algorithms and network layouts such as ConvNets, RNN, LSTMs, Word2Vec, and even the latest transformer architecture
  • The abstraction for the device is perfectly done and its support seamlessly for multiple GPU and even TPU will bring a lot of performance gain for enterprise scoped solution while still keep the flexibility
  • The TensorBoard is amazing. I haven't seen a similar thing in other frameworks on the market. It allows us to quickly understand and debug the model with the info visualization which makes understanding much better
  • A very supportive community, which is the key for sharing the ideas and find the quick and best solutions
Read full review
Cons
  • Memory management. Very weak on that.
  • PySpark not as robust as scala with spark.
  • spark master HA is needed. Not as HA as it should be.
  • Locality should not be a necessity, but does help improvement. But would prefer no locality
Read full review
  • It would be much better if they could provide good documentation and easy ways to understand concepts.
  • It is difficult to understand the concept behind for example, Tensor Graph, which takes a lot of time.
  • As you have to write everything, it is time consuming to write the implementation of whole neural network. It would be better if they can provide some wrapper library to make things easier.
Read full review
Likelihood to Renew
Capacity of computing data in cluster and fast speed.
Read full review
No answers on this topic
Usability
If the team looking to use Apache Spark is not used to debug and tweak settings for jobs to ensure maximum optimizations, it can be frustrating. However, the documentation and the support of the community on the internet can help resolve most issues. Moreover, it is highly configurable and it integrates with different tools (eg: it can be used by dbt core), which increase the scenarios where it can be used
Read full review
Support of multiple components and ease of development.
Read full review
Support Rating
1. It integrates very well with scala or python. 2. It's very easy to understand SQL interoperability. 3. Apache is way faster than the other competitive technologies. 4. The support from the Apache community is very huge for Spark. 5. Execution times are faster as compared to others. 6. There are a large number of forums available for Apache Spark. 7. The code availability for Apache Spark is simpler and easy to gain access to. 8. Many organizations use Apache Spark, so many solutions are available for existing applications.
Read full review
Community support for TensorFlow is great. There's a huge community that truly loves the platform and there are many examples of development in TensorFlow. Often, when a new good technique is published, there will be a TensorFlow implementation not long after. This makes it quick to ally the latest techniques from academia straight to production-grade systems. Tooling around TensorFlow is also good. TensorBoard has been such a useful tool, I can't imagine how hard it would be to debug a deep neural network gone wrong without TensorBoard.
Read full review
Implementation Rating
No answers on this topic
Use of cloud for better execution power is recommended.
Read full review
Alternatives Considered
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only
Read full review
Can't seem to choose any deep learning platform in the above, so I'll list it here: 1. Apache MXNet: this has been used for one of our main algorithms for search as an end-to-end pipeline. We chose this because of the Scala bindings, which makes it easier to integrate with out JVM backend. MXNet seems comparable to TensorFlow, although community support is not as good as TensorFlow, and there are issues with memory leaks that are being worked on. TensorFlow in general is easier to use, but MXNet isn't too far behind. 2. Keras: still a favorite. Often I use this when paired with TensorFlow. TensorFlow 2.0 will make it even easier. 3. PyTorch: only used it a little, so it's hard to provide a good opinion. 4. DL4J: used it initially in an early days project because it has good JVM support. Harder to used not because of poor API design, but because community support is lacking and features don't come out as fast as TensorFlow.
Read full review
Return on Investment
  • Faster turn around on feature development, we have seen a noticeable improvement in our agile development since using Spark.
  • Easy adoption, having multiple departments use the same underlying technology even if the use cases are very different allows for more commonality amongst applications which definitely makes the operations team happy.
  • Performance, we have been able to make some applications run over 20x faster since switching to Spark. This has saved us time, headaches, and operating costs.
Read full review
  • Positive Impact- As I mentioned before its open source. Very easy to learn for average programmer/ developer. We were able to design a POC model for understanding the patient appointment cancellation snd reasons behind it in 3 week time frame.
  • Negative Impact- If you are using tensor flow for small project it works fine. If you are trying to build a model for face recognition it will be hard to program and train the system. It needs data to be processed before hand cannot learn on the go.
Read full review
ScreenShots