Databricks offers the Databricks Lakehouse Platform (formerly the Unified Analytics Platform), a data science platform and Apache Spark cluster manager. The Databricks Unified Data Service provides a platform for data pipelines, data lakes, and data platforms.
$0.07
Per DBU
IBM InfoSphere Information Server
Score 8.0 out of 10
N/A
IBM InfoSphere Information Server is a data integration platform used to understand, cleanse, monitor and transform data. The offerings provide massively parallel processing (MPP) capabilities.
N/A
IBM Watson Studio
Score 10.0 out of 10
N/A
IBM Watson Studio enables users to build, run and manage AI models, and optimize decisions at scale across any cloud. IBM Watson Studio enables users can operationalize AI anywhere as part of IBM Cloud Pak® for Data, the IBM data and AI platform. The vendor states the solution simplifies AI lifecycle management and accelerates time to value with an open, flexible multicloud architecture.
I also use Microsoft Azure Machine Learning in parallel with Databricks. They use different file formats which teach me to be flexible and able to write different programs. They are equally useful to me and I would like to master both platforms for any future usage. I do prefer …
I don't really see other tools that I also use such as R-Studio, eclipse, local Jupyter notebooks, PyCharm, and even Jupyter Labs as being direct competitors to DSx. Mainly because even with things like git integration, more focused on the seem to be on local development as …
Medium to Large data throughput shops will benefit the most from Databricks Spark processing. Smaller use cases may find the barrier to entry a bit too high for casual use cases. Some of the overhead to kicking off a Spark compute job can actually lead to your workloads taking longer, but past a certain point the performance returns cannot be beat.
Information Server is extremely useful to replace manual developments that require a lot of coding effort. It significantly increases the productivity of the initial development and the future maintenance of the processes since it has a visual development environment with self-documentation.
It has a lot of features that are good for teams working on large-scale projects and continuously developing and reiterating their data project models. Really helpful when dealing with large data. It is a kind of one-stop solution for all data science tasks like visualization, cleaning, analyzing data, and developing models but small teams might find a lot of features unuseful.
Because it is an amazing platform for designing experiments and delivering a deep dive analysis that requires execution of highly complex queries, as well as it allows to share the information and insights across the company with their shared workspaces, while keeping it secured.
in terms of graph generation and interaction it could improve their UI and UX
One of the best customer and technology support that I have ever experienced in my career. You pay for what you get and you get the Rolls Royce. It reminds me of the customer support of SAS in the 2000s when the tools were reaching some limits and their engineer wanted to know more about what we were doing, long before "data science" was even a name. Databricks truly embraces the partnership with their customer and help them on any given challenge.
I received answers mostly at once and got answered even further my question: they gave me interesting points of view and suggestion for deepening in the learning path
The most important differentiating factor for Databricks Lakehouse Platform from these other platforms is support for ACID transactions and the time travel feature. Also, native integration with managed MLflow is a plus. EMR, Cloudera, and Hortonworks are not as optimized when it comes to Spark Job Execution. Other platforms need to be self-managed, which is another huge hassle.
The main reason I personally changed over from Azure ML Studio is because it lacked any support for significant custom modelling with packages and services such as TensorFlow, scikit-learn, Microsoft Cognitive Toolkit and Spark ML. IBM Watson Studio provides these services and does so in a well integrated and easy to use fashion making it a preferable service over the other services that I have personally used.