TrustRadius: an HG Insights company

Datadog

Score8.8 out of 10

371 Reviews and Ratings

What is Datadog?

Datadog is a monitoring service for IT, Dev and Ops teams who write and run applications at scale, and want to turn the massive amounts of data produced by their apps, tools and services into actionable insight.

Read more details.

Media

Screenshot of the out-of-the-box and customizable monitoring dashboards.
Screenshot of Datadog's collaboration features, where users can discuss issues in-context with production data, annotate changes and notify their teams, see who responded to that alert before, and discover what was done to fix it.
Screenshot of where Datadog unifies traces, metrics, and logs—the three pillars of observability.
Screenshot of some of Datadog's 400+ built-in integrations.
Screenshot of Datadog's Service Map, which decomposes an application into all its component services and draws the observed dependencies between these services in real time
Screenshot of centralized log data, pulled from any source.
Screenshot of Datadog's Host Map, which lets users see all hosts together on one screen, grouped and filtered as desired, with metrics made instantly comprehensible via color and shape.

1 / 7

Screenshot of the out-of-the-box and customizable monitoring dashboards.

Top Performing Features

  • Data visualization

    Ability to produce visualizations to simplify understanding of data

    Category average: 8.8

  • Remote monitoring

    Monitoring of network operational activities through the use of remote devices known as monitors or probes

    Category average: 8.9

  • Automated alerts and notifications

    System generates alerts and notifications to provide timely intervention if required

    Category average: 9

Areas for Improvement

  • Administrator access control

    Management and restriction of user access to specific system assets

    Category average: 6.4

  • Software and hardware inventory

    Management of the physical and software from acquisition through disposal.

    Category average: 6.2

  • Policy-based automation

    Policy-based management is an administrative approach used to simplify the system management by drafting rules to deal with common situations

    Category average: 7.2

Who Buys & Uses Datadog

Pros

  • Powerful data visualization and customizable dashboards for comprehensive insights
  • Effective log management and search capabilities for efficient analysis
  • Intelligent and easily configurable alerting mechanisms

Cons

  • Opaque and potentially high pricing model, especially with increased data ingestion
  • Steep learning curve due to complex query syntax and extensive features
  • Challenges with dashboard customization and navigating the user interface

Datadog Robust and Feature-Rich Observability with Excellent Monitoring Capabilities.

Use Cases and Deployment Scope

In my company, the customer is using our application, which is highly dependent on API calls. We use Datadog to gain maximum visibility into the customer experience, including all API calls and RUM recordings, to understand what customers are facing. In addition, our QA and support teams are highly dependent on Datadog to identify what the customer did wrong and what gaps we have in our application.

Pros

  • Logs
  • RUM recordings
  • Metrics
  • Failures

Cons

  • Payload in API calls.
  • Sometimes the logs are missing.

Return on Investment

  • We spend less time in identifying the problems using Datadog.
  • We are getting great from Datadog by allocating less resources on investigations.

Usability

Alternatives Considered

Amazon CloudWatch

Other Software Used

Amazon CloudWatch, AWS Lambda, Amazon S3 (Simple Storage Service)

Top notch product with cutting edge features

Use Cases and Deployment Scope

Datadog serves as our primary observability platform and helps us to maintain responsive incident management strategy for our product. Datadog covers many use cases which includes generating key custom metrics for top-level analytics of our data, APM traces that tracks fine-grained microservices communication with continuous log aggregation and multiple test suites for our product.

Pros

  • Log aggregation with seamless search indexes
  • APM traces that describes each call to every microservice and their communication
  • All out-of-the box and third party integrations within Datadog
  • Keeping with the AI trends with LLM monitoring assistance

Cons

  • Datadog on-call service

Return on Investment

  • Our platform includes complex multi-tenant, multi-cloud environment. Datadog's core features such as log monitoring, microservices trace generation and custom test suites definition enhanced our platform's performance and availability

Usability

Other Software Used

Prometheus, Grafana Loki

Why we choose Datadog over other services.

Use Cases and Deployment Scope

We use Datadog to assess logs and rum sessions & review the scope of our code. It's a fascinating tool, to be very honest; we switched over to it after using other services. Datadog is superior to its competitors, offering a one-stop shop for all your needs. Now, we're used to the platform and can't move forward without it.

Pros

  • Rum sessions.
  • Error logs.
  • Client preference views.

Cons

  • Firebase scoop.
  • Parallel automation testing.
  • Local pipeline runs.

Return on Investment

  • Better cost.
  • Multiple resources.
  • Great log management.

Usability

Alternatives Considered

Shipbook, AWS CloudTrail and GitLab

Other Software Used

GitLab, Shipbook, Firebase Crashlytics

Datadog user experience supporting a banking solution

Use Cases and Deployment Scope

With the Platform Reliability Engineering team supporting legacy core banking applications that has multiple services was always a challenge. With Datadog APM distributed tracing, anomaly detection algorithms for the transactions break, supporting the infrastructure monitoring for both onprem and kubernetes cluster is made easy from our org. Monitoring the core banking application is our usecase.

Pros

  • APM - particularly with Dynamic instrumentation helpful for trace analysis
  • Infrastructure Monitoring- for both on-premise - host map and OCP deployments from Kubernetes explorer
  • BIT AI SRE Agent - for Incident investigations, finding RCA

Cons

  • Support for traceability of mainframes application, applications/solutions on C/C++
  • Avoiding duplicating/non useful monitors and allow AI monitoring to decide, which should be mandatory monitors
  • Custom metrics usability to have similar kind of visualizations of various metrics

Return on Investment

  • Incident response time drastically decreased
  • Reporting is made easy with the dashboards and single places for all important widgets
  • Though with some false positives anomaly detection monitoring was helpful

Usability

Alternatives Considered

Splunk AppDynamics, Prometheus and Grafana

Other Software Used

Finacle Core Banking, Red Hat JBoss Enterprise Application Platform, Red Hat Ansible Automation Platform

Excellent Tool in the Modern Observability Toolbox

Use Cases and Deployment Scope

We use Datadog to monitor so many different tools across our whole tech stack. For my use case we have integrations with so many different platforms and Datadog helps us to gain visibility into failures to help with troubleshooting and keep a pulse on overall performance. I use the bits tool at times to use natural language to get a summary of data trends. Often we will look at the payloads, export the logs to csv for further analysis. But lately I've found that bits does a good job of summarizing so I don't have to export as much.

Pros

  • Surfaces failures in my DMS/ERP integrations
  • Summarizes issues with the bits AI tool
  • Helps me to see exactly what payload values are problematic

Cons

  • I know there are ways to set up alerts when something passes a threshold but I can't seem to figure it out

Return on Investment

  • It has had a huge impact in helping increase CSAT by surfacing issues before they become problems. We can ofter get ahead of issues before things get heated.

Usability

Other Software Used

Atlassian Jira, Looker Studio, Amazon Athena