Likelihood to Recommend One of AWS Glue's most notable features that aid in the creation and transformation of data is its data catalog. Support, scheduling, and the automation of the data schema recognition make it superior to its competitors aside from that. It also integrates perfectly with other AWS tools. The main restriction may be integrated with systems outside of the AWS environment. It functions flawlessly with the current AWS services but not with other goods. Another potential restriction that comes to mind is that glue operates on a spark, which means the engineer needs to be conversant in the language.
Read full review Paxata can be highly useful to someone who doesn't like/have any experience with writing codes to treat data before using it as input into BI dashboards. Paxata can accelerate data cleaning in environments where a large amount of unclean data is generated and business decisions on the go are required. It performs really well while dealing with natural language.
Read full review Pros It is extremely fast, easy, and self-intuitive. Though it is a suite of services, it requires pretty less time to get control over it. As it is a managed service, one need not take care of a lot of underlying details. The identification of data schema, code generation, customization, and orchestration of the different job components allows the developers to focus on the core business problem without worrying about infrastructure issues. It is a pay-as-you-go service. So, there is no need to provide any capacity in advance. So, it makes scheduling much easier. Read full review Visualize distributions in large data sets effectively which enable the user to quickly spot outliers and treat them appropriately Provides recommendation to merge datasets based on matching column values The cluster and edit feature in my opinion is its most powerful feature and reduces cardinality in column with text Read full review Cons In-Stream schema registries feature people can not use this more efficiently in Connections feature they can add more connectors as well The crucial problem with AWS Glue is that it only works with AWS. Read full review Doesn't provide recommendation on how to impute values There is a lag quite often We can say whether a column has errors or quality issues in the first look Read full review Support Rating Amazon responds in good time once the ticket has been generated but needs to generate tickets frequent because very few sample codes are available, and it's not cover all the scenarios.
Read full review Alternatives Considered AWS Glue is a fully managed ETL service that automates many ETL tasks, making it easier to set AWS Glue simplifies ETL through a visual interface and automated code generation.
Read full review Paxata is a much better tool when it comes to handling natural language but
Talend provides recommendations on how to impute missing values and outliers. Paxata provides recommendations on dataset tie-ups and joins but
Talend doesn't provide any such recommendations. In paxata you can visualize distribution of data in a column and filter them by dragging and selecting the section you'd like to retain
Read full review Return on Investment It had a positive impact on the way we build our data lake. It is the single source of truth for data structure (schemas/tables/views). Read full review It saves time to clean data It reduces the requirement of too many data engineer/stewards and hence adds positive impact on the return of the business Read full review ScreenShots