Apache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.
N/A
Snowflake
Score 8.7 out of 10
N/A
The Snowflake Cloud Data Platform is the eponymous data warehouse with, from the company in San Mateo, a cloud and SQL based DW that aims to allow users to unify, integrate, analyze, and share previously siloed data in secure, governed, and compliant ways. With it, users can securely access the Data Cloud to share live data with customers and business partners, and connect with other organizations doing business as data consumers, data providers, and data service providers.
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only.
Apache Spark is a fast-processing in-memory computing framework. It is 10 times faster than Apache Hadoop. Earlier we were using Apache Hadoop for processing data on the disk but now we are shifted to Apache Spark because of its in-memory computation capability. Also in SAP …
There are a few alternatives that can do the same transformation and aggregation like Apache Spark can do but most of them are not able to perform parallel computation. For example, pandas is a really good tool to do that but not parallelized; However, there are some tools that …
Apache Spark has much more better performance and features if we compare with Hive or map/reduce kind of solutions. Spark has many other features for machine learning, streaming.
1. Apache Spark is almost 100 % faster than Hadoop. 2. Apache Spark is more stable than Amazon EMR. 3. The end to end distributed machine library is more robust in Apache Spark.
Databricks uses Spark as a foundation, and is also a great platform. It does bring several add-ons, which we did not feel needed by the time we evaluated - and haven't needed since then. One interesting plus in our opinion was the engineering support, which is great depending …
It is easy to learn, read and to maintain. It brings the best of the Ruby on Rails framework from Java that helps to create a web service so easily. Communication is one of the most distinctive features of Apache Spark compared to alternative products. You are able to …
We evaluated SAS alongside with Apache Spark but during the course of proof of concept found that Apache Spark was able to support the hadoop eco-system and hadoop file system much better. It was much faster at that time while having the ability to process data quickly for the …
Consultor Tecnico - Java Developer and Php Developer.
Chose Apache Spark
I prefer Apache Spark compared to Hadoop, since in my experience Spark has more usability and comes equipped with simple APIs for Scala, Python, Java and Spark SQL, as well as provides feedback in REPL format on the commands. At the same time, Apache Spark seems to have the …
All the above systems work quite well on big data transformations whereas Spark really shines with its bigger API support and its ability to read from and write to multiple data sources. Using Spark one can easily switch between declarative versus imperative versus functional …
Even with Python, MapReduce is lengthy coding. Combination of Python with Apache Spark will not only shorten the code, but it will effectively increase the speed of algorithms. Occasionally, I use MapReduce, but Apache Spark will replace MapReduce very soon. It has many …
vs MapRedce, it was faster and easier to manage. Especially for Machine Learning, where MapReduce is lacking. Also Apache Storm was slower and didn't scale as much as Spark does. Spark elasticity was easier to apply compared to storm and MapReduce. managing resources for …
Spark in comparison to similar technologies ends up being a one stop shop. You can achieve so much with this one framework instead of having to stitch and weave multiple technologies from the Hadoop stack, all while getting incredibility performance, minimal boilerplate, and …
Apache Pig and Apache Hive provide most of the things spark provide but apache spark has more features like actions and transformations which are easy to code. Spark uses optimization technique as we can select driver program and manipulate DAG (Directed Acyclic Graph) Python …
There are a few newer frameworks for general processing like Flink, Beam, frameworks for streaming like Samza and Storm, and traditional Map-Reduce. I think Spark is at a sweet spot where its clearly better than Map-Reduce for many workflows yet has gotten a good amount of …
Spark has primarily replaced my use of writing pure Hadoop MapReduce or Apache Pig jobs for processing data. I like the fact that I can alternate between the main programming languages that I know - Java and Python - and use those to learn the Scala API. Spark also can be …
We use these tools for applications they are better suited for vs a Snowflake. For e.g. MS Fabric has powerful agentic AI capabilities; Redshift is our go to choice for the TMT vertical within the organization and Databricks is the default choice for AI/ML applications.
Snowflake provides various features, such as integration with Python using Snowpark. The reporting feature that caters to your small reporting needs is Snowsight. The Snowflake data marketplace is where you can get multiple data for free and even some of the data which you can …
These are comparable products that can make sense depending on the specific needs of your organization. All are certainly serviceable and have varying pros and cons. Snowflake seems to provide the greatest degree of flexibility and easy scalability as new data gets brought into …
We needed scalability and a new way of organizing our data; Snowflake allowed us to have a clearer view of our data warehouses and schemas. Snowflake is also way superior in terms of speed and quick insights from the raw data you query, which is very valuable to us.
Snowflake has an attractive pricing model with auto-suspend and auto-resume and pay per use. AWS Redshift requires higher administrative efforts to maintain and scale the platform whereas with Snowflake those admin tasks are not needed or automatically taken care of.
We had a MS SQL server with over 2 TB of ram & 51 processors that we were using, that could no longer handle our workload. Snowflake can handle 3 times that workload with ease and efficiency.
Snowflake is much faster and easier to write queries and pull data. But the visualization part of Snowflake is not as good as them. Also, Snowflake only supports SQL queries but not python or other languages. So basically Snowflake is the expert in its field but not suitable …
We particularly liked Snowflake's security model as well as its unique storage (whereby everything is essentially a pointer to immutable micro-partitions, which is the key behind its zero-copy cloning, its secure sharing, its time travel, etc.). and also how it separates …
While Snowflake is more open to cloud eco system, SAP integrated well with SAP eco system products like SAP ECC or SAP S/4. So for people who have invested heavily in SAP eco system including SAP ECC or S/4, it makes sense to go with SAP DWC which is also evolving very rapidly. …
In my opinion, the other tools have similar and some different features; however, when I ran proof of technologies between Synapse and Snowflake. Snowflake did things better or just had functionality that the other tools did not. One that stuck out at the time was scale up …
Each of the other solutions were cloud vendor specific, Snowflake can ride on either Amazon Web Services, Microsoft Azure, or Google Cloud. The fact that they are ANSI-sql compliant and have an effective means of offloading data makes them portable and easy to sell to teams …
Azure and Snowflake compared very similarly, but Snowflake provided more options to integrate and connect with tools/companies that were not partners. It seemed to be a more flexible environment. The barrier for entry on Oracle and Google we just too complicated. In particular, …
I have had the experience of using one more database management system at my previous workplace. What Snowflake provides is better user-friendly consoles, suggestions while writing a query, ease of access to connect to various BI platforms to analyze, [and a] more robust system …
Snowflake has won the match because it is giving an excellent performance with its efficient features and reliable results. This is a totally secure program for our precious and important data.
Our initial data warehousing solution was Treasure Data. We had issues with the costly pricing model, which would be exhorbitant if we want to hold our data in memory and query using Presto. As a result, some heavy lifting was done in Hive (managed by Treasure Data); …
In my experience running the data management practice at InterWorks, we believe that cloud data warehouse products will eventually serve the majority of data warehousing use cases and power data analytics at most companies. Of this cohort, we believe that Snowflake is the best …
Redshift compute and storage can be scaled up/down together (though they added some features recently, they don't quite add up). I haven't tried Avalanche or Firebolt but would love to in the near future, due to their pedigree or revolutionary billing methods.
- Cost was the main aspect on the decision. - Performance was in par or better compared to other tools in the market. - Snowflake in my opinion stacks better than other tools I have used in the past.
Accommodates future data types such as JSON and XML. Scalability is another advantage. Pay per use is beneficial for organizations like yours. Direct connectors with AWS help us to go with it. No limit on user creation and clone data not eating up extra disk space are a few …
Since we switch from amazon redshift to Snowflake, we found Snowflake is much better than redshift in many ways, including the data integrate and data pull. However, comparing directly pull data from amazon s3, Snowflake is quite slow in terms of data pull speed and the more …
Compared to Amazon Redshift, Snowflake is slightly easier and faster to achieve ROI but based on the user's perspective, the two tools have very little difference since both are leveraging SQL to pull data from AWS S3. Snowflake is also working with Microsoft Azure but it is …
Our issue with Redshift was that it was very expensive. On top of that, queries were still slow and if we used more of Redshift's memory, then it would have cost even more. Snowflake is not cheap, but less costly for us. Plus, the performance was much better. Also, we got to …
Apache Spark has rich APIs for regular data transformations or for ML workloads or for graph workloads, whereas other systems may not such a wide range of support. Choose it when you need to perform data transformations for big data as offline jobs, whereas use MongoDB-like distributed database systems for more realtime queries.
If you need a quick query, snowflake is the way to go. It's super simple and scalable; we were struggling before with Azure, and with Snowflake, everything runs smoothly, and we have more control over our schemas and warehouses. Snowflake, in my opinion, is the next step when you want to scale your business and manage data. If your company is still small, there may be cheaper options.
It performs a conventional disk-based process when the data sets are too large to fit into memory, which is very useful because, regardless of the size of the data, it is always possible to store them.
It has great speed and ability to join multiple types of databases and run different types of analysis applications. This functionality is super useful as it reduces work times
Apache Spark uses the data storage model of Hadoop and can be integrated with other big data frameworks such as HBase, MongoDB, and Cassandra. This is very useful because it is compatible with multiple frameworks that the company has, and thus allows us to unify all the processes.
Snowflake scales appropriately allowing you to manage expense for peak and off peak times for pulling and data retrieval and data centric processing jobs
Snowflake offers a marketplace solution that allows you to sell and subscribe to different data sources
Snowflake manages concurrency better in our trials than other premium competitors
Snowflake has little to no setup and ramp up time
Snowflake offers online training for various employee types
Do not force customers to renew for same or higher amount to avoid loosing unused credits. Already paid credits should not expire (at least within a reasonable time frame), independent of renewal deal size.
SnowFlake is very cost effective and we also like the fact we can stop, start and spin up additional processing engines as we need to. We also like the fact that it's easy to connect our SQL IDEs to Snowflake and write our queries in the environment that we are used to
If the team looking to use Apache Spark is not used to debug and tweak settings for jobs to ensure maximum optimizations, it can be frustrating. However, the documentation and the support of the community on the internet can help resolve most issues. Moreover, it is highly configurable and it integrates with different tools (eg: it can be used by dbt core), which increase the scenarios where it can be used
The interface is similar to other SQL query systems I've used and is fairly easy to use. My only complaint is the syntax issues. Another thing is that the error messages are not always the easiest thing to understand, especially when you incorporate temp tables. Some of that is to be expected with any new database.
1. It integrates very well with scala or python. 2. It's very easy to understand SQL interoperability. 3. Apache is way faster than the other competitive technologies. 4. The support from the Apache community is very huge for Spark. 5. Execution times are faster as compared to others. 6. There are a large number of forums available for Apache Spark. 7. The code availability for Apache Spark is simpler and easy to gain access to. 8. Many organizations use Apache Spark, so many solutions are available for existing applications.
We have had terrific experiences with Snowflake support. They have drilled into queries and given us tremendous detail and helpful answers. In one case they even figured out how a particular product was interacting with Snowflake, via its queries, and gave us detail to go back to that product's vendor because the Snowflake support team identified a fault in its operation. We got it solved without lots of back-and-forth or finger-pointing because the Snowflake team gave such detailed information.
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only
Snowflake provides various features, such as integration with Python using Snowpark. The reporting feature that caters to your small reporting needs is Snowsight. The Snowflake data marketplace is where you can get multiple data for free and even some of the data which you can buy according to your needs. And the integration options with various tools like Sigma are add-ons.
Faster turn around on feature development, we have seen a noticeable improvement in our agile development since using Spark.
Easy adoption, having multiple departments use the same underlying technology even if the use cases are very different allows for more commonality amongst applications which definitely makes the operations team happy.
Performance, we have been able to make some applications run over 20x faster since switching to Spark. This has saved us time, headaches, and operating costs.