Apache Spark vs. SQL Server Integration Services (SSIS)

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Spark
Score 8.9 out of 10
N/A
Apache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.N/A
SSIS
Score 8.0 out of 10
N/A
Microsoft's SQL Server Integration Services (SSIS) is a data integration solution.N/A
Pricing
Apache SparkSQL Server Integration Services (SSIS)
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Apache SparkSSIS
Free Trial
NoNo
Free/Freemium Version
NoNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache SparkSQL Server Integration Services (SSIS)
Considered Both Products
Apache Spark
Chose Apache Spark
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only.
Chose Apache Spark
Apache Spark is a fast-processing in-memory computing framework. It is 10 times faster than Apache Hadoop. Earlier we were using Apache Hadoop for processing data on the disk but now we are shifted to Apache Spark because of its in-memory computation capability. Also in SAP …
Chose Apache Spark
Other teams used to work on Apache Hadoop but our team started with Apache Spark directly.
Chose Apache Spark
There are a few alternatives that can do the same transformation and aggregation like Apache Spark can do but most of them are not able to perform parallel computation. For example, pandas is a really good tool to do that but not parallelized; However, there are some tools that …
Chose Apache Spark
  • Apache Spark works in distributed mode using cluster
  • Informatica and Datastage cannot scale horizontally
  • We can write custom code in spark, whereas in Datastage and Informatica we can only choose the different features proivided already.
Chose Apache Spark
Apache Spark has much more better performance and features if we compare with Hive or map/reduce kind of solutions. Spark has many other features for machine learning, streaming.
Chose Apache Spark
Spark is simply awesome to work on with any data sets and also has an in-memory database which makes it very flexible.
Chose Apache Spark
1. Apache Spark is almost 100 % faster than Hadoop.
2. Apache Spark is more stable than Amazon EMR.
3. The end to end distributed machine library is more robust in Apache Spark.
Chose Apache Spark
Databricks uses Spark as a foundation, and is also a great platform. It does bring several add-ons, which we did not feel needed by the time we evaluated - and haven't needed since then. One interesting plus in our opinion was the engineering support, which is great depending …
Chose Apache Spark
It is easy to learn, read and to maintain. It brings the best of the Ruby on Rails framework from Java that helps to create a web service so easily. Communication is one of the most distinctive features of Apache Spark compared to alternative products. You are able to …
Chose Apache Spark
We evaluated SAS alongside with Apache Spark but during the course of proof of concept found that Apache Spark was able to support the hadoop eco-system and hadoop file system much better. It was much faster at that time while having the ability to process data quickly for the …
Chose Apache Spark
I prefer Apache Spark compared to Hadoop, since in my experience Spark has more usability and comes equipped with simple APIs for Scala, Python, Java and Spark SQL, as well as provides feedback in REPL format on the commands. At the same time, Apache Spark seems to have the …
Chose Apache Spark
All the above systems work quite well on big data transformations whereas Spark really shines with its bigger API support and its ability to read from and write to multiple data sources. Using Spark one can easily switch between declarative versus imperative versus functional …
Chose Apache Spark
Even with Python, MapReduce is lengthy coding. Combination of Python with Apache Spark will not only shorten the code, but it will effectively increase the speed of algorithms. Occasionally, I use MapReduce, but Apache Spark will replace MapReduce very soon. It has many …
Chose Apache Spark
vs MapRedce, it was faster and easier to manage. Especially for Machine Learning, where MapReduce is lacking. Also Apache Storm was slower and didn't scale as much as Spark does. Spark elasticity was easier to apply compared to storm and MapReduce.
managing resources for …
Chose Apache Spark
We specifically choose Spark over MapReduce to make the cluster processing faster
Chose Apache Spark
Spark in comparison to similar technologies ends up being a one stop shop. You can achieve so much with this one framework instead of having to stitch and weave multiple technologies from the Hadoop stack, all while getting incredibility performance, minimal boilerplate, and …
Chose Apache Spark
Apache Pig and Apache Hive provide most of the things spark provide but apache spark has more features like actions and transformations which are easy to code. Spark uses optimization technique as we can select driver program and manipulate DAG (Directed Acyclic Graph)
Python …
Chose Apache Spark
There are a few newer frameworks for general processing like Flink, Beam, frameworks for streaming like Samza and Storm, and traditional Map-Reduce. I think Spark is at a sweet spot where its clearly better than Map-Reduce for many workflows yet has gotten a good amount of …
Chose Apache Spark
Spark has primarily replaced my use of writing pure Hadoop MapReduce or Apache Pig jobs for processing data. I like the fact that I can alternate between the main programming languages that I know - Java and Python - and use those to learn the Scala API. Spark also can be …
SSIS
Chose SSIS
Informatica PowerCenter (legacy) and CloverDX
Chose SSIS
Both are very similar. Azure is cloud based. It is easier for the organization who uses cloud based application. The SQL Server Integration Services is cost effective. Azure was more on the expensive side for our organization. Azure was a little complex, it needed special …
Chose SSIS
I think SQL Server Integration Services is better suited for on-premises data movement and ADF is more suited for the cloud. Though ADF has more connectors, SQL Server Integration Services is more robust and has better functionality just because it has been around much longer
Chose SSIS
Fivetran, Stitch, and Etleap are all 1000x more modern than SSIS and 100x less aggravating. While those tools are mainly used to sync data rather than transform it, the ELT model works much better than the ETL model in most situations.
Chose SSIS
We just selected SSIS because we use SQL Server Management System (SSMS) to manage our database. As SSIS is a component of the Microsoft SQL Server there are no problems with integration and everything works perfectly. In addition, we don't have to learn how to use another …
Chose SSIS
Low-cost relative to other products - in fact, zero cost if one is considering the license cost as being for the database engine with Integration Services added on. It has a comparable range of functionality and performance and as such it's a 'no-brainer' to use SSIS over …
Chose SSIS
SnapLogic and Azure Data Factory are better than SQL Server Integration Services mostly because they are Integration Platform as a Service (IPAAS) services, whereas SQL Server Integration Services is an on-premise. So the basic differences such as, need a VPN to connect to the …
Chose SSIS
SSIS is similar to Alteryx and Informatica PowerCenter in a way because these are all drag-and-drop ETL tools with similar functionality. Alteryx is a step ahead because it has some advanced ETL functionalities including statistical calculations etc. and a better ability to set …
Chose SSIS
Alteryx Designer is easier to use for machine learning models. The functionality of drag and drop is the most valuable. It is a very user-friendly tool that can be understood easily. My teams also work with other solutions, such as Integration Services, and these solutions are …
Chose SSIS
I had nothing to do with the choice or install. I assume it was made because it's easy to integrate with our SQL Server environment and free. I'm not sure of any other enterprise level solution that would solve this problem, but I would likely have approached it with …
Chose SSIS
SQL Server Integration Services is a good alternative to cut down on costs and have more flexibility on developing.
Chose SSIS
SAP Business Objects was a primary concurrent software against the MS SSIS but it has a more steep learning curve and requires additional investment into the SAP-related software infrastructure. With SSIS one can start easily with simple data extraction / DTL tools of Express …
Chose SSIS
I personally prefer SSIS. There are items that each do better than the others, but the ease of use of SSIS, along with its extensibility to 3rd party, ability to write any code required in the tool, and uses the same IDE for the MS BI suite (more of an issue if you're not a …
Chose SSIS
I used the Pentaho Data Integration (PDI) ETL tool. The PDI ETL tool does not have a public user collection like the SQL Server Integration Services(SSIS) ETL tool. Therefore, you may not be able to find instant solutions for your problems. But it has advantages over the SSIS …
Chose SSIS
We selected SSIS as it came part of our standard SQL license. We did not evaluate any other solutions as SSIS has met all our needs.
Chose SSIS
SQL Server is already in our wheelhouse so it only made sense to utilize the tools we already had available to us--SSIS, SSAS, & SSRS. Other non-technical users seem to be more comfortable using alternatives to SSIS. However, these alternatives are not as good as SSIS at …
Chose SSIS
These are all great products and, honestly, can move data faster. They include more enterprise features and have some great qualities about each. However, they all cost a lot depending on the implementation you need. With SQL Server Integration Services, you do not have any …
Chose SSIS
When looking to evaluate different options, we looked first to the experience and software we had in-house that would accomplish the job. When assessing alternatives outside we were looking for the tool that would offer the most flexibility.

SSIS provided the most robust set of …
Chose SSIS
It’s basically a free tool and it has more features than anyone would ever need. If you look online for answers for SISS packages you will find a world of information that can cover almost any situation for your business. This tool can be used in any business and it provides …
Chose SSIS
SSIS is a very basic, developer-oriented ETL tool and while it lacks many of the nice UX features of its competitors it is a powerful tool that comes as a part of SQL Server and, in the hands of experienced developers with domain knowledge, can meet most organizations' ETL …
Chose SSIS
SSIS and Denodo differ in their approaches to ETL and Data integrations. SSIS is more affordable from a cost and licensing perspective (if you have Microsoft licensing), but Denodo is no slouch. If you go with Denodo, you are not creating data, there are pros and cons to …
Chose SSIS
SQL Server Integration Services does a good job for our SQL Server environments and was selected for that reason. For a SQL Server-only implementations, I would recommend SQL Server Integration Services. When we compared SSIS to other ETL providers against SQL Server, SSIS was …
Features
Apache SparkSQL Server Integration Services (SSIS)
Data Source Connection
Comparison of Data Source Connection features of Product A and Product B
Apache Spark
-
Ratings
SQL Server Integration Services (SSIS)
7.0
Ratings
17% below category average
Connect to traditional data sources00 Ratings9.00 Ratings
Connecto to Big Data and NoSQL00 Ratings5.00 Ratings
Data Transformations
Comparison of Data Transformations features of Product A and Product B
Apache Spark
-
Ratings
SQL Server Integration Services (SSIS)
6.8
Ratings
17% below category average
Simple transformations00 Ratings9.00 Ratings
Complex transformations00 Ratings4.70 Ratings
Data Modeling
Comparison of Data Modeling features of Product A and Product B
Apache Spark
-
Ratings
SQL Server Integration Services (SSIS)
7.5
Ratings
5% below category average
Data model creation00 Ratings9.00 Ratings
Metadata management00 Ratings6.00 Ratings
Business rules and workflow00 Ratings7.00 Ratings
Collaboration00 Ratings9.00 Ratings
Testing and debugging00 Ratings6.30 Ratings
Data Governance
Comparison of Data Governance features of Product A and Product B
Apache Spark
-
Ratings
SQL Server Integration Services (SSIS)
5.3
Ratings
41% below category average
Integration with data quality tools00 Ratings6.10 Ratings
Integration with MDM tools00 Ratings4.60 Ratings
Best Alternatives
Apache SparkSQL Server Integration Services (SSIS)
Small Businesses

No answers on this topic

Skyvia
Skyvia
Score 10.0 out of 10
Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 8.6 out of 10
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache SparkSQL Server Integration Services (SSIS)
Likelihood to Recommend
9.0
(0 ratings)
8.0
(0 ratings)
Likelihood to Renew
10.0
(0 ratings)
9.0
(0 ratings)
Usability
8.0
(0 ratings)
8.0
(0 ratings)
Performance
-
(0 ratings)
8.8
(0 ratings)
Support Rating
8.7
(0 ratings)
8.0
(0 ratings)
Implementation Rating
-
(0 ratings)
10.0
(0 ratings)
User Testimonials
Apache SparkSQL Server Integration Services (SSIS)
Likelihood to Recommend
Apache Spark has rich APIs for regular data transformations or for ML workloads or for graph workloads, whereas other systems may not such a wide range of support. Choose it when you need to perform data transformations for big data as offline jobs, whereas use MongoDB-like distributed database systems for more realtime queries.
Read full review
Ideal for daily standard ETL use cases whether the data is sourced from / transferred to the native connectors (like SQL Server) or FTP. Best if the company uses MS suite of tools. There are better options in the market for chaining tasks where you want a custom flow of executions depending on the outcome of each process or if you want advanced functionality like API connections, etc.
Read full review
Pros
  • It performs a conventional disk-based process when the data sets are too large to fit into memory, which is very useful because, regardless of the size of the data, it is always possible to store them.
  • It has great speed and ability to join multiple types of databases and run different types of analysis applications. This functionality is super useful as it reduces work times
  • Apache Spark uses the data storage model of Hadoop and can be integrated with other big data frameworks such as HBase, MongoDB, and Cassandra. This is very useful because it is compatible with multiple frameworks that the company has, and thus allows us to unify all the processes.
Read full review
  • SSIS works very well pulling well-defined data into SQL Server from a wide variety of data sources.
  • It comes free with the SQL Server so it is hard not to consider using it providing you have a team who is trained and experienced using SSIS.
  • When SSIS doesn't have exactly what you need you can use C# or VBA to extend its functionality.
Read full review
Cons
  • Memory management. Very weak on that.
  • PySpark not as robust as scala with spark.
  • spark master HA is needed. Not as HA as it should be.
  • Locality should not be a necessity, but does help improvement. But would prefer no locality
Read full review
  • SSIS memory usage can be quite high particularly when SSI and SQL server are on the same machine
  • SSIS is not available on any environment other than Microsoft Windows
  • SSIS does not function with any database engine back-end other than Microsoft SQL Server
Read full review
Likelihood to Renew
Capacity of computing data in cluster and fast speed.
Read full review
SSIS is responsible for running core business processed managing core business data. It can be managed, improved and expanded using minimal internal resources. It is also able to support all of our current data infrastructure. Replacing SSIS would be time consuming and costly with no apparent ROI.
Read full review
Usability
If the team looking to use Apache Spark is not used to debug and tweak settings for jobs to ensure maximum optimizations, it can be frustrating. However, the documentation and the support of the community on the internet can help resolve most issues. Moreover, it is highly configurable and it integrates with different tools (eg: it can be used by dbt core), which increase the scenarios where it can be used
Read full review
It is easy to learn, works great on many features, but needs improvement on ETL troubleshooting and performance monitoring functionality. Great tool on Microsoft stack. it is great with simple, structured datasets. Once logic gets fancy like nested conditionals
complex joins,
reusable transformations, versioned logic …SQL Server Integration Services (SSIS) packages become hard to read and harder to maintain. Source control is painful. Errors can be cryptic
Logging takes effort to set up well
Debugging in production is limited.
Read full review
Performance
No answers on this topic
Raw performance is great. At times, depending on the machine you are using for development, the IDE can have issues. Deploying projects is very easy and the tool set they give you to monitor jobs out of the box is decent. If you do very much with it you will have to write into your projects performance tracking though.
Read full review
Support Rating
1. It integrates very well with scala or python. 2. It's very easy to understand SQL interoperability. 3. Apache is way faster than the other competitive technologies. 4. The support from the Apache community is very huge for Spark. 5. Execution times are faster as compared to others. 6. There are a large number of forums available for Apache Spark. 7. The code availability for Apache Spark is simpler and easy to gain access to. 8. Many organizations use Apache Spark, so many solutions are available for existing applications.
Read full review
The support, when necessary, is excellent. But beyond that, it is very rarely necessary because the user community is so large, vibrant and knowledgable, a simple Google query or forum question can answer almost everything you want to know. You can also get prewritten script tasks with a variety of functionality that saves a lot of time.
Read full review
Implementation Rating
No answers on this topic
The implementation may be different in each case, it is important to properly analyze all the existing infrastructure to understand the kind of work needed, the type of software used and the compatibility between these, the features that you want to exploit, to understand what is possible and which ones require integration with third-party tools
Read full review
Alternatives Considered
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only
Read full review
Both are very similar. Azure is cloud based. It is easier for the organization who uses cloud based application. The SQL Server Integration Services is cost effective. Azure was more on the expensive side for our organization. Azure was a little complex, it needed special training to use it. Azure was not accurate with complex data.
Read full review
Return on Investment
  • Faster turn around on feature development, we have seen a noticeable improvement in our agile development since using Spark.
  • Easy adoption, having multiple departments use the same underlying technology even if the use cases are very different allows for more commonality amongst applications which definitely makes the operations team happy.
  • Performance, we have been able to make some applications run over 20x faster since switching to Spark. This has saved us time, headaches, and operating costs.
Read full review
  • Without this, we would have to manually update a spreadsheet of our SQL Server inventory
  • We would also have poor alerting; if an instance was down we wouldn't know until it was reported by a user
  • We only have one other person who uses SQL Server Integration Services , he's the expert. It would fall to me without him and I would not enjoy being responsible for it.
Read full review
ScreenShots