Apache Sqoop vs. Databricks Data Intelligence Platform

Apache Sqoop

Apache Sqoop

4 Reviews and Ratings

Databricks Data Intelligence Platform

Databricks Data Intelligence Platform

108 Reviews and Ratings

Overview
Product	Rating	Most Used By	Product Summary	Starting Price
Apache Sqoop	Score 8.8 out of 10	N/A	Apache Sqoop is a tool for use with Hadoop, used to transfer data between Apache Hadoop and other, structured data stores.	N/A
Databricks Data Intelligence Platform	Score 8.8 out of 10	N/A	Databricks offers the Databricks Lakehouse Platform (formerly the Unified Analytics Platform), a data science platform and Apache Spark cluster manager. The Databricks Unified Data Service provides a platform for data pipelines, data lakes, and data platforms.	$0.07 Per DBU

Pricing

Apache Sqoop

Databricks Data Intelligence Platform

Editions & Modules

No answers on this topic

Standard: $0.07
Per DBU
Premium: $0.10
Per DBU
Enterprise: $0.13
Per DBU

Offerings

Pricing Offerings
Apache Sqoop	Databricks Data Intelligence Platform
Free Trial
No	No
Free/Freemium Version
No	No
Premium Consulting/Integration Services
No	No

Entry-level Setup Fee

No setup fee

No setup fee

Additional Details

—

—

More Pricing Information

Community Pulse
	Apache Sqoop	Databricks Data Intelligence Platform

Best Alternatives
	Apache Sqoop	Databricks Data Intelligence Platform
Small Businesses	No answers on this topic	No answers on this topic
Medium-sized Companies	Cloudera Manager Score 9.9 out of 10	Snowflake Score 8.7 out of 10
Enterprises	IBM Analytics Engine Score 7.1 out of 10	Snowflake Score 8.7 out of 10
All Alternatives	View all alternatives	View all alternatives

User Ratings
	Apache Sqoop	Databricks Data Intelligence Platform
Likelihood to Recommend	9.0 (1 ratings)	9.4 (21 ratings)
Usability	- (0 ratings)	9.7 (7 ratings)
Support Rating	- (0 ratings)	8.7 (2 ratings)
Contract Terms and Pricing Model	- (0 ratings)	8.0 (1 ratings)
Professional Services	- (0 ratings)	10.0 (1 ratings)

User Testimonials
	Apache Sqoop	Databricks Data Intelligence Platform
Likelihood to Recommend	Apache Sqoop is great for sending data between a JDBC compliant database and a Hadoop environment. Sqoop is built for those who need a few simple CLI options to import a selection of database tables into Hadoop, do large dataset analysis that could not commonly be done with that database system due to resource constraints, then export the results back into that database (or another). Sqoop falls short when there needs to be some extra, customized processing between database extract, and Hadoop loading, in which case Apache Spark's JDBC utilities might be preferred Incentivized Jordan Moore Consultant Read full review	Databricks Medium to Large data throughput shops will benefit the most from Databricks Spark processing. Smaller use cases may find the barrier to entry a bit too high for casual use cases. Some of the overhead to kicking off a Spark compute job can actually lead to your workloads taking longer, but past a certain point the performance returns cannot be beat. Incentivized Austin Franchino Senior Data and Security Engineer Read full review
Pros	Apache Provides generalized JDBC extensions to migrate data between most database systems Generates Java classes upon reading database records for use in other code utilizing Hadoop's client libraries Allows for both import and export features Incentivized Jordan Moore Consultant Read full review	Databricks Process raw data in One Lake (S3) env to relational tables and views Share notebooks with our business analysts so that they can use the queries and generate value out of the data Try out PySpark and Spark SQL queries on raw data before using them in our Spark jobs Modern day ETL operations made easy using Databricks. Provide access mechanism for different set of customers Incentivized Verified User Anonymous Read full review
Cons	Apache Sqoop2 development seems to have stalled. I have set it up outside of a Cloudera CDH installation, and I actually prefer it's "Sqoop Server" model better than just the CLI client version that is Sqoop1. This works especially well in a microservices environment, where there would be only one place to maintain the JDBC drivers to use for Sqoop. Incentivized Jordan Moore Consultant Read full review	Databricks Sometimes, when multiple jobs depend on each other in different environments, it is not always easy to see the full workflow in one place. It is sometimes difficult to determine which job or cluster contributes more to the overall cost. For beginners, cluster configuration may be a little difficult. So more recommendation in the platform can help. Incentivized Verified User Anonymous Read full review
Usability	Apache No answers on this topic	Databricks Because it is an amazing platform for designing experiments and delivering a deep dive analysis that requires execution of highly complex queries, as well as it allows to share the information and insights across the company with their shared workspaces, while keeping it secured. in terms of graph generation and interaction it could improve their UI and UX Incentivized Verified User Anonymous Read full review
Support Rating	Apache No answers on this topic	Databricks One of the best customer and technology support that I have ever experienced in my career. You pay for what you get and you get the Rolls Royce. It reminds me of the customer support of SAS in the 2000s when the tools were reaching some limits and their engineer wanted to know more about what we were doing, long before "data science" was even a name. Databricks truly embraces the partnership with their customer and help them on any given challenge. Jonatan Bouchard Director Data Science Read full review
Alternatives Considered	Apache Sqoop comes preinstalled on the major Hadoop vendor distributions as the recommended product to import data from relational databases. The ability to extend it with additional JDBC drivers makes it very flexible for the environment it is installed within. Spark also has a useful JDBC reader, and can manipulate data in more ways than Sqoop, and also upload to many other systems than just Hadoop. Kafka Connect JDBC is more for streaming database updates using tools such as Oracle GoldenGate or Debezium. Streamsets and Apache NiFi both provide a more "flow based programming" approach to graphically laying out connectors between various systems, including JDBC and Hadoop. Incentivized Jordan Moore Consultant Read full review	Databricks The most important differentiating factor for Databricks Lakehouse Platform from these other platforms is support for ACID transactions and the time travel feature. Also, native integration with managed MLflow is a plus. EMR, Cloudera, and Hortonworks are not as optimized when it comes to Spark Job Execution. Other platforms need to be self-managed, which is another huge hassle. Incentivized Verified User Anonymous Read full review
Return on Investment	Apache When combined with Cloudera's HUE, it can enable non-technical users to easily import relational data into Hadoop. Being able to manipulate large datasets in Hadoop, and them load them into a type of "materialized view" in an external database system has yielded great insights into the Hadoop datalake without continuously running large batch jobs. Sqoop isn't very user-friendly for those uncomfortable with a CLI. Incentivized Jordan Moore Consultant Read full review	Databricks The ability to spin up a BIG Data platform with little infrastructure overhead allows us to focus on business value not admin DB has the ability to terminate/time out instances which helps manage cost. The ability to quickly access typical hard to build data scenarios easily is a strength. Incentivized Verified User Anonymous Read full review
ScreenShots