Apache Hive vs. Posit vs. Presto

Apache Hive

95 Reviews and Ratings

Posit

236 Reviews and Ratings

Presto

19 Reviews and Ratings

Overview
Product	Rating	Most Used By	Product Summary	Starting Price
Apache Hive	Score 8.0 out of 10	N/A	Apache Hive is database/data warehouse software that supports data querying and analysis of large datasets stored in the Hadoop distributed file system (HDFS) and other compatible systems, and is distributed under an open source license.	N/A
Posit	Score 10.0 out of 10	N/A	Posit, formerly RStudio, is a modular data science platform, combining open source and commercial products.	N/A
Presto	Score 10.0 out of 10	N/A	Presto is an open source SQL query engine designed to run queries on data stored in Hadoop or in traditional databases. Teradata supported development of Presto followed the acquisition of Hadapt and Revelytix.	N/A

Pricing

Apache Hive

Posit

Presto

Editions & Modules

No answers on this topic

Offerings

Pricing Offerings
Apache Hive	Posit	Presto
Free Trial
No	Yes	No
Free/Freemium Version
No	Yes	No
Premium Consulting/Integration Services
No	No	No

Entry-level Setup Fee

No setup fee

Optional

No setup fee

Additional Details

—

More Pricing Information

Community Pulse
	Apache Hive	Posit	Presto
Considered Multiple Products	Apache Hive Verified User Analyst Chose Apache Hive Presto is slightly less reliable but much faster for interactive querying. These tools would not be replacements for each other, but rather complements. Incentivized Helpful? Praveen Murugesan Engineering Manager - Ride Experience Chose Apache Hive We selected Hive because it supports SQL, schema and provides structure on top of hadoop. Having data structured has its benefits, especially if there are thousands of users processing on the same data over and over again. Pig provides the ability to process unstructured data. … Incentivized Helpful? Ananth Gouri Assistant Professor Chose Apache Hive One of the major advantages of using Presto or the main reason why people use Presto (Teradata) is due to that fact it can support multiple data sources - which is lacking as in the case of Apache Hive. But still, most people who come from a Structured data-based background … Incentivized Helpful? Verified User C-Level Executive Chose Apache Hive Community support and ease of use -not deployment. It enables querying and analyzing large amounts of data stored in HDFS, on the petabyte scale. It has a query language called HQL that transforms SQL queries into MapReduce jobs that run on Hadoop, and it is wonderful for the … Incentivized Helpful? Verified User Administrator Chose Apache Hive Due to effective queries resolved time and the performance and user-friendly framework compared to other products. Incentivized Helpful? Jordan Moore Staff Consultant Chose Apache Hive Hive was one of the first SQL on Hadoop technologies, and it comes bundled with the main Hadoop distributions of HDP and CDH. Since its release, it has gained good improvements, but selecting the right SQL on Hadoop technology requires a good understanding of the strengths and … Incentivized Helpful?	Posit Verified User Manager Chose Posit RStudio's user interface is easier to use than Jupyter Notebook (particularly for users that are new to programming). Many of our users have experience with RStudio Desktop, so switching to RStudio Server Pro was very easy. Deploying applications is also much easier thanks to … Incentivized Helpful?	Presto Praveen Murugesan Engineering Manager - Ride Experience Chose Presto I think Presto is one of the best solutions out there today at the cutting edge for interactive query analysis. One of the challenges is presto is a niche tool for the interactive query use case and doesn't have the knobs and whistles as much as Spark. In the foreseeable future … Incentivized Helpful?

Features

Apache Hive

Posit

Presto

Platform Connectivity

Comparison of Platform Connectivity features of Product A and Product B
	Apache Hive - Ratings	Posit 9.3 27 Ratings 11% above category average	Presto - Ratings
Connect to Multiple Data Sources	00 Ratings	8.026 Ratings	00 Ratings
Extend Existing Data Sources	00 Ratings	9.927 Ratings	00 Ratings
Automatic Data Format Detection	00 Ratings	9.926 Ratings	00 Ratings

Data Exploration

Comparison of Data Exploration features of Product A and Product B
	Apache Hive - Ratings	Posit 9.0 27 Ratings 6% above category average	Presto - Ratings
Visualization	00 Ratings	8.027 Ratings	00 Ratings
Interactive Data Analysis	00 Ratings	10.024 Ratings	00 Ratings

Data Preparation

Comparison of Data Preparation features of Product A and Product B
	Apache Hive - Ratings	Posit 10.0 26 Ratings 20% above category average	Presto - Ratings
Interactive Data Cleaning and Enrichment	00 Ratings	10.024 Ratings	00 Ratings
Data Transformations	00 Ratings	10.026 Ratings	00 Ratings

Platform Data Modeling

Comparison of Platform Data Modeling features of Product A and Product B
	Apache Hive - Ratings	Posit 10.0 22 Ratings 17% above category average	Presto - Ratings
Multiple Model Development Languages and Tools	00 Ratings	10.022 Ratings	00 Ratings
Single platform for multiple model development	00 Ratings	10.022 Ratings	00 Ratings
Self-Service Model Delivery	00 Ratings	10.019 Ratings	00 Ratings

Model Deployment

Comparison of Model Deployment features of Product A and Product B
	Apache Hive - Ratings	Posit 9.9 18 Ratings 15% above category average	Presto - Ratings
Flexible Model Publishing Options	00 Ratings	10.018 Ratings	00 Ratings
Security, Governance, and Cost Controls	00 Ratings	9.915 Ratings	00 Ratings

Best Alternatives
	Apache Hive	Posit	Presto
Small Businesses	Google BigQuery Score 8.8 out of 10	Jupyter Notebook Score 8.5 out of 10	InterSystems IRIS Score 8.0 out of 10
Medium-sized Companies	Cloudera Enterprise Data Hub Score 9.0 out of 10	Mathematica Score 7.0 out of 10	InterSystems IRIS Score 8.0 out of 10
Enterprises	Oracle Exadata Score 9.8 out of 10	Dataiku Score 8.5 out of 10	SAP IQ Score 10.0 out of 10
All Alternatives	View all alternatives	View all alternatives	View all alternatives

User Ratings
	Apache Hive	Posit	Presto
Likelihood to Recommend	8.0 (35 ratings)	10.0 (123 ratings)	7.8 (2 ratings)
Likelihood to Renew	10.0 (1 ratings)	9.7 (17 ratings)	- (0 ratings)
Usability	8.5 (7 ratings)	8.0 (4 ratings)	- (0 ratings)
Availability	- (0 ratings)	9.4 (3 ratings)	- (0 ratings)
Support Rating	7.0 (6 ratings)	8.9 (9 ratings)	- (0 ratings)
Implementation Rating	- (0 ratings)	9.3 (4 ratings)	- (0 ratings)
Configurability	- (0 ratings)	10.0 (1 ratings)	- (0 ratings)
Product Scalability	- (0 ratings)	8.2 (3 ratings)	- (0 ratings)

User Testimonials
	Apache Hive	Posit	Presto
Likelihood to Recommend	Apache Software work execution is on a large scale, it is good to use for new projects or organizational changes, data lineage mapping has always been dubious but this one has had good results. You can store and synchronize data from different departments, the storage process can be manual but it is best automated. Incentivized Camilo Palacios Administrador informático. Read full review	Posit (formerly RStudio) In my humble opinion, if you are working on something related to Statistics, RStudio is your go-to tool. But if you are looking for something in Machine Learning, look out for Python. The beauty is that there are packages now by which you can write Python/SQL in R. Cross-platform functionality like such makes RStudio way ahead of its competition. A couple of chinks in RStudio armor are very small and can be considered as nagging just for the sake of argument. Other than completely based on programming language, I couldn't find significant drawbacks to using RStudio. It is one of the best free software available in the market at present. Incentivized Verified User Anonymous Read full review	Open Source Presto is for interactive simple queries, where Hive is for reliable processing. If you have a fact-dim join, presto is great..however for fact-fact joins presto is not the solution.. Presto is a great replacement for proprietary technology like Vertica Incentivized Praveen Murugesan Engineering Manager - Ride Experience Read full review
Pros	Apache Apache Hive allows use to write expressive solutions to complex problems thanks to its SQL-like syntax. Relatively easy to set up and start using. Very little ramp-up to start using the actual product, documentation is very thorough, there is an active community, and the code base is constantly being improved. Incentivized Verified User Anonymous Read full review	Posit (formerly RStudio) The support is incredibly professional and helpful, and they often go out of their way to help me when something doesn't work. The one-click publishing from RStudio Connect is absolutely amazing, and I really like the way that it deploys your exact package versions, because otherwise, you can get in a terrible mess. Python doesn't feel quite as native as R at the moment but I have definitely deployed stuff in R and Python that works beautifully which is really nice indeed. Incentivized Chris Beeley Senior Analyst Read full review	Open Source Linking, embedding links and adding images is easy enough. Once you have become familiar with the interface, Presto becomes very quick & easy to use (but, you have to practice & repeat to know what you are doing - it is not as intuitive as one would hope). Organizing & design is fairly simple with click & drag parameters. Incentivized Corinne Nacin-Martinez Consumer Sales & Sevice Manager Read full review
Cons	Apache Some queries, particularly complex joins, are still quite slow and can take hours Previous jobs and queries are not stored sometimes Switching to Impala can sometimes be time-consuming (i.e. the system hangs, or is slow to respond). Sometimes, directories and tables don't load properly which causes confusion Incentivized Verified User Anonymous Read full review	Posit (formerly RStudio) Python integration is newer and still can be rough, especially with when using virtual environments. RStudio Connect pricing feels very department focused, not quite an enterprise perspective. Some of the RStudio packages don't follow conventional development guidelines (API breaking changes with minor version numbers) which can make supporting larger projects over longer timeframes difficult. Incentivized B. Mark Ewing Digital Corporate Technology Manager – R&D Read full review	Open Source Presto was not designed for large fact fact joins. This is by design as presto does not leverage disk and used memory for processing which in turn makes it fast.. However, this is a tradeoff..in an ideal world, people would like to use one system for all their use cases, and presto should get exhaustive by solving this problem. Resource allocation is not similar to YARN and presto has a priority queue based query resource allocation..so a query that takes long takes longer...this might be alleviated by giving some more control back to the user to define priority/override. UDF Support is not available in presto. You will have to write your own functions..while this is good for performance, it comes at a huge overhead of building exclusively for presto and not being interoperable with other systems like Hive, SparkSQL etc. Incentivized Praveen Murugesan Engineering Manager - Ride Experience Read full review
Likelihood to Renew	Apache Since I do not know the second data warehouse solution that integrate with HDFS as well as Hive. Yinghua Hu Senior Data Scientist Read full review	Posit (formerly RStudio) There is no viable alternative right now. The toolset is good and the functionality is increasing with every release. It is backed by regular releases and ongoing development by the RStudio team. There is good engagement with RStudio directly when support is required. Also there's a strong and growing community of developers who provide additional support and sample code. Incentivized Verified User Anonymous Read full review	Open Source No answers on this topic
Usability	Apache Hive is a very good big data analysis and ad-hoc query platform, which supports scaling also. The BI processes can be easily integrated with Hadoop via the Hive. It can deal with a much larger data set that traditional RDBMS can not. It is a "must-have" component of the big data domain. Incentivized Verified User Anonymous Read full review	Posit (formerly RStudio) For someone who learns how to use the software and picks up on the "language" of R, it's very easy to use. For beginners, it can be hard and might require a course, as well as the appropriate statistical training to understand what packages to use and when Incentivized Verified User Anonymous Read full review	Open Source No answers on this topic
Reliability and Availability	Apache No answers on this topic	Posit (formerly RStudio) RStudio is very available and cheap to use. It needs to be updated every once in a while, but the updates tend to be quick and they do not hinder my ability to make progress. I have not experienced any RStudio outages, and I have used the application quite a bit for a variety of statistical analyses Incentivized Kenton Woods Graduate Research Assistant Read full review	Open Source No answers on this topic
Support Rating	Apache Apache Hive is a FOSS project and its open source. We need not definitely comment on anything about the support of open source and its developer community. But, it has got tremendous developer support, awesome documentation. I would justify the fact that much support can be gathered from the community backup. Incentivized Ananth Gouri Assistant Professor Read full review	Posit (formerly RStudio) Since R is trendy among statisticians, you can find lots of help from the data science/ stats communities. If you need help with anything related to RStudio or R, google it or search on StackOverflow, you might easily find the solution that you are looking for. Incentivized Xiaotong Song Business Analyst Read full review	Open Source No answers on this topic
Implementation Rating	Apache No answers on this topic	Posit (formerly RStudio) We did it at the individual level: anyone willing to code in R can use it. No real deployment involved. Verified User Anonymous Read full review	Open Source No answers on this topic
Alternatives Considered	Apache Besides Hive, I have used Google BigQuery, which is costly but have very high computation speed. Amazon Redshift is the another product, I used in my recent organisation. Both Redshift and BigQuery are managed solution whereas Hive needs to be managed Incentivized Manjeet Singh Senior Manager - Engineering Read full review	Posit (formerly RStudio) RStudio was provided as the most customizable. It was also strictly the most feature-rich as far as enabling our organization to script, run, and make use of R open-source packages in our data analysis workstreams. It also provided some support for python, which was useful when we had R heavy code with some python threaded in. Overall we picked Rstudio for the features it provided for our data analysis needs and the ability to interface with our existing resources. Incentivized Verified User Anonymous Read full review	Open Source Presto is good for a templated design appeal. You cannot be too creative via this interface - but, the layout and options make the finalized visual product appealing to customers. The other design products I use are for different purposes and not really comparable to Presto. Incentivized Corinne Nacin-Martinez Consumer Sales & Sevice Manager Read full review
Scalability	Apache No answers on this topic	Posit (formerly RStudio) RStudio is very scalable as a product. The issue I have is that it doesn't necessarily fit in nicely with the mainly Microsoft environment that everybody else is using. Having RStudio for us means dedicated servers and recruiting staff who know how to manage the environment. This isn't a fault of the product at all, it's just part of the data science landscape that we all have to put up with. Having said that RStudio is absolutely great for running on low spec servers and there are loads of options to handle concurrency, memory use, etc. Incentivized Chris Beeley Senior Analyst Read full review	Open Source No answers on this topic
Return on Investment	Apache Apache hive is secured and scalable solution that helps in increasing the overall organization productivity. Apache hive can handle and process large amount of data in a sufficient time manner. It simplifies writing SQL queries, hence helping the organization as most companies use SQL for all query jobs. Incentivized Verified User Anonymous Read full review	Posit (formerly RStudio) Using it for data science in a very big and old company, the most positive impact, from my point of view, has been the ability of spreading data culture across the group. Shortening the path from data to value. Still it's hard to quantify economic benefits, we are struggling and it's a great point of attention, since splitting out the contribution of the single aspects of a project (and getting the RStudio pie) is complicated. What is sure is that, in the long run, RStudio is boosting productivity and making the process in which is embedded more efficient (cost reduction). Incentivized Flavio Leccese Group Data Scientist Read full review	Open Source Presto has helped scale Uber's interactive data needs. We have migrated a lot out of proprietary tech like Vertica. Presto has helped build data driven applications on its stack than maintain a separate online/offline stack. Presto has helped us build data exploration tools by leveraging it's power of interactive and is immensely valuable for data scientists. Incentivized Praveen Murugesan Engineering Manager - Ride Experience Read full review
ScreenShots		Posit Screenshots