Apache Spark vs. Snowflake

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Spark
Score 8.9 out of 10
N/A
Apache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.N/A
Snowflake
Score 8.7 out of 10
N/A
The Snowflake Cloud Data Platform is the eponymous data warehouse with, from the company in San Mateo, a cloud and SQL based DW that aims to allow users to unify, integrate, analyze, and share previously siloed data in secure, governed, and compliant ways. With it, users can securely access the Data Cloud to share live data with customers and business partners, and connect with other organizations doing business as data consumers, data providers, and data service providers.N/A
Pricing
Apache SparkSnowflake
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Apache SparkSnowflake
Free Trial
NoYes
Free/Freemium Version
NoNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache SparkSnowflake
Considered Both Products
Apache Spark
Chose Apache Spark
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only.
Chose Apache Spark
Apache Spark is a fast-processing in-memory computing framework. It is 10 times faster than Apache Hadoop. Earlier we were using Apache Hadoop for processing data on the disk but now we are shifted to Apache Spark because of its in-memory computation capability. Also in SAP …
Chose Apache Spark
Other teams used to work on Apache Hadoop but our team started with Apache Spark directly.
Chose Apache Spark
There are a few alternatives that can do the same transformation and aggregation like Apache Spark can do but most of them are not able to perform parallel computation. For example, pandas is a really good tool to do that but not parallelized; However, there are some tools that …
Chose Apache Spark
  • Apache Spark works in distributed mode using cluster
  • Informatica and Datastage cannot scale horizontally
  • We can write custom code in spark, whereas in Datastage and Informatica we can only choose the different features proivided already.
Chose Apache Spark
Apache Spark has much more better performance and features if we compare with Hive or map/reduce kind of solutions. Spark has many other features for machine learning, streaming.
Chose Apache Spark
Spark is simply awesome to work on with any data sets and also has an in-memory database which makes it very flexible.
Chose Apache Spark
1. Apache Spark is almost 100 % faster than Hadoop.
2. Apache Spark is more stable than Amazon EMR.
3. The end to end distributed machine library is more robust in Apache Spark.
Chose Apache Spark
Databricks uses Spark as a foundation, and is also a great platform. It does bring several add-ons, which we did not feel needed by the time we evaluated - and haven't needed since then. One interesting plus in our opinion was the engineering support, which is great depending …
Chose Apache Spark
It is easy to learn, read and to maintain. It brings the best of the Ruby on Rails framework from Java that helps to create a web service so easily. Communication is one of the most distinctive features of Apache Spark compared to alternative products. You are able to …
Chose Apache Spark
We evaluated SAS alongside with Apache Spark but during the course of proof of concept found that Apache Spark was able to support the hadoop eco-system and hadoop file system much better. It was much faster at that time while having the ability to process data quickly for the …
Chose Apache Spark
I prefer Apache Spark compared to Hadoop, since in my experience Spark has more usability and comes equipped with simple APIs for Scala, Python, Java and Spark SQL, as well as provides feedback in REPL format on the commands. At the same time, Apache Spark seems to have the …
Chose Apache Spark
All the above systems work quite well on big data transformations whereas Spark really shines with its bigger API support and its ability to read from and write to multiple data sources. Using Spark one can easily switch between declarative versus imperative versus functional …
Chose Apache Spark
Even with Python, MapReduce is lengthy coding. Combination of Python with Apache Spark will not only shorten the code, but it will effectively increase the speed of algorithms. Occasionally, I use MapReduce, but Apache Spark will replace MapReduce very soon. It has many …
Chose Apache Spark
vs MapRedce, it was faster and easier to manage. Especially for Machine Learning, where MapReduce is lacking. Also Apache Storm was slower and didn't scale as much as Spark does. Spark elasticity was easier to apply compared to storm and MapReduce.
managing resources for …
Chose Apache Spark
We specifically choose Spark over MapReduce to make the cluster processing faster
Chose Apache Spark
Spark in comparison to similar technologies ends up being a one stop shop. You can achieve so much with this one framework instead of having to stitch and weave multiple technologies from the Hadoop stack, all while getting incredibility performance, minimal boilerplate, and …
Chose Apache Spark
Apache Pig and Apache Hive provide most of the things spark provide but apache spark has more features like actions and transformations which are easy to code. Spark uses optimization technique as we can select driver program and manipulate DAG (Directed Acyclic Graph)
Python …
Chose Apache Spark
There are a few newer frameworks for general processing like Flink, Beam, frameworks for streaming like Samza and Storm, and traditional Map-Reduce. I think Spark is at a sweet spot where its clearly better than Map-Reduce for many workflows yet has gotten a good amount of …
Chose Apache Spark
Spark has primarily replaced my use of writing pure Hadoop MapReduce or Apache Pig jobs for processing data. I like the fact that I can alternate between the main programming languages that I know - Java and Python - and use those to learn the Scala API. Spark also can be …
Snowflake
Chose Snowflake
We use these tools for applications they are better suited for vs a Snowflake. For e.g. MS Fabric has powerful agentic AI capabilities; Redshift is our go to choice for the TMT vertical within the organization and Databricks is the default choice for AI/ML applications.
Chose Snowflake
Snowflake provides various features, such as integration with Python using Snowpark. The reporting feature that caters to your small reporting needs is Snowsight. The Snowflake data marketplace is where you can get multiple data for free and even some of the data which you can …
Chose Snowflake
These are comparable products that can make sense depending on the specific needs of your organization. All are certainly serviceable and have varying pros and cons. Snowflake seems to provide the greatest degree of flexibility and easy scalability as new data gets brought into …
Chose Snowflake
Snowflake stacks up well against many of the other players in the space. We used Teradata before and it seems to perform better.
Chose Snowflake
We needed scalability and a new way of organizing our data; Snowflake allowed us to have a clearer view of our data warehouses and schemas. Snowflake is also way superior in terms of speed and quick insights from the raw data you query, which is very valuable to us.
Chose Snowflake
Snowflake has an attractive pricing model with auto-suspend and auto-resume and pay per use. AWS Redshift requires higher administrative efforts to maintain and scale the platform whereas with Snowflake those admin tasks are not needed or automatically taken care of.
Chose Snowflake
Cloud -Native architecture, Separation of Storage and Compute provides flexibility, Scalable Compute, Cost and Data Security
Chose Snowflake
We had a MS SQL server with over 2 TB of ram & 51 processors that we were using, that could no longer handle our workload. Snowflake can handle 3 times that workload with ease and efficiency.
Chose Snowflake
Snowflake is much faster and easier to write queries and pull data. But the visualization part of Snowflake is not as good as them. Also, Snowflake only supports SQL queries but not python or other languages. So basically Snowflake is the expert in its field but not suitable …
Chose Snowflake
We particularly liked Snowflake's security model as well as its unique storage (whereby everything is essentially a pointer to immutable micro-partitions, which is the key behind its zero-copy cloning, its secure sharing, its time travel, etc.). and also how it separates …
Chose Snowflake
While Snowflake is more open to cloud eco system, SAP integrated well with SAP eco system products like SAP ECC or SAP S/4. So for people who have invested heavily in SAP eco system including SAP ECC or S/4, it makes sense to go with SAP DWC which is also evolving very rapidly. …
Chose Snowflake
In my opinion, the other tools have similar and some different features; however, when I ran proof of technologies between Synapse and Snowflake. Snowflake did things better or just had functionality that the other tools did not. One that stuck out at the time was scale up …
Chose Snowflake
Each of the other solutions were cloud vendor specific, Snowflake can ride on either Amazon Web Services, Microsoft Azure, or Google Cloud. The fact that they are ANSI-sql compliant and have an effective means of offloading data makes them portable and easy to sell to teams …
Chose Snowflake
Azure and Snowflake compared very similarly, but Snowflake provided more options to integrate and connect with tools/companies that were not partners. It seemed to be a more flexible environment. The barrier for entry on Oracle and Google we just too complicated. In particular, …
Chose Snowflake
I have had the experience of using one more database management system at my previous workplace. What Snowflake provides is better user-friendly consoles, suggestions while writing a query, ease of access to connect to various BI platforms to analyze, [and a] more robust system …
Chose Snowflake
Snowflake has won the match because it is giving an excellent performance with its efficient features and reliable results. This is a totally secure program for our precious and important data.
Chose Snowflake
Our initial data warehousing solution was Treasure Data. We had issues with the costly pricing model, which would be exhorbitant if we want to hold our data in memory and query using Presto. As a result, some heavy lifting was done in Hive (managed by Treasure Data); …
Chose Snowflake
Snowflake beats these other products in every category it was rated against
Chose Snowflake
In my experience running the data management practice at InterWorks, we believe that cloud data warehouse products will eventually serve the majority of data warehousing use cases and power data analytics at most companies. Of this cohort, we believe that Snowflake is the best …
Chose Snowflake
Redshift compute and storage can be scaled up/down together (though they added some features recently, they don't quite add up). I haven't tried Avalanche or Firebolt but would love to in the near future, due to their pedigree or revolutionary billing methods.
Chose Snowflake
- Cost was the main aspect on the decision.
- Performance was in par or better compared to other tools in the market.
- Snowflake in my opinion stacks better than other tools I have used in the past.
Chose Snowflake
Accommodates future data types such as JSON and XML. Scalability is another advantage. Pay per use is beneficial for organizations like yours. Direct connectors with AWS help us to go with it. No limit on user creation and clone data not eating up extra disk space are a few …
Chose Snowflake
Since we switch from amazon redshift to Snowflake, we found Snowflake is much better than redshift in many ways, including the data integrate and data pull. However, comparing directly pull data from amazon s3, Snowflake is quite slow in terms of data pull speed and the more …
Chose Snowflake
Compared to Amazon Redshift, Snowflake is slightly easier and faster to achieve ROI but based on the user's perspective, the two tools have very little difference since both are leveraging SQL to pull data from AWS S3. Snowflake is also working with Microsoft Azure but it is …
Chose Snowflake
Our issue with Redshift was that it was very expensive. On top of that, queries were still slow and if we used more of Redshift's memory, then it would have cost even more. Snowflake is not cheap, but less costly for us. Plus, the performance was much better. Also, we got to …
Best Alternatives
Apache SparkSnowflake
Small Businesses

No answers on this topic

Google BigQuery
Google BigQuery
Score 8.7 out of 10
Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
Google BigQuery
Google BigQuery
Score 8.7 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 8.6 out of 10
Google BigQuery
Google BigQuery
Score 8.7 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache SparkSnowflake
Likelihood to Recommend
9.0
(0 ratings)
8.8
(0 ratings)
Likelihood to Renew
10.0
(0 ratings)
7.0
(0 ratings)
Usability
8.0
(0 ratings)
8.9
(0 ratings)
Support Rating
8.7
(0 ratings)
9.9
(0 ratings)
User Testimonials
Apache SparkSnowflake
Likelihood to Recommend
Apache Spark has rich APIs for regular data transformations or for ML workloads or for graph workloads, whereas other systems may not such a wide range of support. Choose it when you need to perform data transformations for big data as offline jobs, whereas use MongoDB-like distributed database systems for more realtime queries.
Read full review
If you need a quick query, snowflake is the way to go. It's super simple and scalable; we were struggling before with Azure, and with Snowflake, everything runs smoothly, and we have more control over our schemas and warehouses. Snowflake, in my opinion, is the next step when you want to scale your business and manage data. If your company is still small, there may be cheaper options.
Read full review
Pros
  • It performs a conventional disk-based process when the data sets are too large to fit into memory, which is very useful because, regardless of the size of the data, it is always possible to store them.
  • It has great speed and ability to join multiple types of databases and run different types of analysis applications. This functionality is super useful as it reduces work times
  • Apache Spark uses the data storage model of Hadoop and can be integrated with other big data frameworks such as HBase, MongoDB, and Cassandra. This is very useful because it is compatible with multiple frameworks that the company has, and thus allows us to unify all the processes.
Read full review
  • Snowflake scales appropriately allowing you to manage expense for peak and off peak times for pulling and data retrieval and data centric processing jobs
  • Snowflake offers a marketplace solution that allows you to sell and subscribe to different data sources
  • Snowflake manages concurrency better in our trials than other premium competitors
  • Snowflake has little to no setup and ramp up time
  • Snowflake offers online training for various employee types
Read full review
Cons
  • Memory management. Very weak on that.
  • PySpark not as robust as scala with spark.
  • spark master HA is needed. Not as HA as it should be.
  • Locality should not be a necessity, but does help improvement. But would prefer no locality
Read full review
  • Add constraints for views and not just for tables
  • Do not force customers to renew for same or higher amount to avoid loosing unused credits. Already paid credits should not expire (at least within a reasonable time frame), independent of renewal deal size.
Read full review
Likelihood to Renew
Capacity of computing data in cluster and fast speed.
Read full review
SnowFlake is very cost effective and we also like the fact we can stop, start and spin up additional processing engines as we need to. We also like the fact that it's easy to connect our SQL IDEs to Snowflake and write our queries in the environment that we are used to
Read full review
Usability
If the team looking to use Apache Spark is not used to debug and tweak settings for jobs to ensure maximum optimizations, it can be frustrating. However, the documentation and the support of the community on the internet can help resolve most issues. Moreover, it is highly configurable and it integrates with different tools (eg: it can be used by dbt core), which increase the scenarios where it can be used
Read full review
The interface is similar to other SQL query systems I've used and is fairly easy to use. My only complaint is the syntax issues. Another thing is that the error messages are not always the easiest thing to understand, especially when you incorporate temp tables. Some of that is to be expected with any new database.
Read full review
Support Rating
1. It integrates very well with scala or python. 2. It's very easy to understand SQL interoperability. 3. Apache is way faster than the other competitive technologies. 4. The support from the Apache community is very huge for Spark. 5. Execution times are faster as compared to others. 6. There are a large number of forums available for Apache Spark. 7. The code availability for Apache Spark is simpler and easy to gain access to. 8. Many organizations use Apache Spark, so many solutions are available for existing applications.
Read full review
We have had terrific experiences with Snowflake support. They have drilled into queries and given us tremendous detail and helpful answers. In one case they even figured out how a particular product was interacting with Snowflake, via its queries, and gave us detail to go back to that product's vendor because the Snowflake support team identified a fault in its operation. We got it solved without lots of back-and-forth or finger-pointing because the Snowflake team gave such detailed information.
Read full review
Alternatives Considered
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only
Read full review
Snowflake provides various features, such as integration with Python using Snowpark. The reporting feature that caters to your small reporting needs is Snowsight. The Snowflake data marketplace is where you can get multiple data for free and even some of the data which you can buy according to your needs. And the integration options with various tools like Sigma are add-ons.
Read full review
Return on Investment
  • Faster turn around on feature development, we have seen a noticeable improvement in our agile development since using Spark.
  • Easy adoption, having multiple departments use the same underlying technology even if the use cases are very different allows for more commonality amongst applications which definitely makes the operations team happy.
  • Performance, we have been able to make some applications run over 20x faster since switching to Spark. This has saved us time, headaches, and operating costs.
Read full review
  • With separate compute and storage feature, the queries get executed quickly and it improves our overall productivity.
  • Earlier we were using a different product for analytical purposes, but with Snowflake's in-built analytical feature we are now able to save money.
  • Snowflake is cost efficient, features like auto suspend for compute resources helped to control the costs.
Read full review
ScreenShots

Snowflake Screenshots

Screenshot of Snowflake Installation