Nancy Zhu
Product Manager
Evan Marcantonio
Senior Product Manager
Nicole Parisi
Product Marketing Manager
Chris Miller
Staff Engineer
When an issue in production triggers an alert, the people responding to it are often working in Slack while the evidence they need is elsewhere. Responders need to move between conversations, telemetry data, source code, and incident tooling as they form hypotheses, coordinate actions, and keep stakeholders informed. That context switching can slow down a time-sensitive investigation and make updates harder to follow.
Bits Chat brings Datadog’s natural-language interface into Slack, giving responders access to Bits AI from the channel where they’re already collaborating. During an incident, teams can ask Bits to investigate the alert, pull in relevant telemetry data, generate a pull request for the fix, and keep track of follow-ups without leaving Slack.
In this post, we’ll follow an incident from alert to resolution and show how you can manage your incidents directly in Slack to:
Start an investigation
Let’s say you’re an on-call engineer for an ecommerce website and you receive a monitor notification that the website’s recommendation service is experiencing errors and failing because of timeouts. When you declare an incident, Datadog creates a dedicated Slack channel so that you and your team can collaborate on fixing the issue. Instead of leaving the conversation to begin gathering evidence, you can type @Datadog investigate
in the channel to ask Bits to start an investigation.
Bits Investigation analyzes the issue by forming hypotheses from relevant telemetry data, runbooks, and past incidents. You can also include additional context or a hypothesis in the initial @Datadog investigate
message to help Bits focus its investigation from the start. As the investigation runs, Bits posts updates in Slack so that responders can follow its work alongside their own discussion.
Troubleshoot with Bits and your responders
When Bits Investigation finishes, it returns its root cause findings and recommended next steps to the Slack conversation. In our example of the ecommerce website, Bits identifies a recent code change as the likely cause of the issue and shares the supporting telemetry data.
Team members can then mention @Datadog
to ask questions about the findings, go deeper into analyzing the telemetry data, compare the findings with what the team has observed, and decide how to resolve the issue. For example, they can ask which endpoints and customer regions are affected and whether the errors are affecting any downstream services. All of these activities can happen in the same Slack channel, and the investigation stays connected to the related discussion.
Act on the investigation
After responders identify the likely cause, the next challenge is turning that finding into action. Bits Remediation suggests next steps based on the investigation and enables teams to take actions directly from the incident conversation in Slack. Those actions can include adding more responders, running incident workflows, or posting Status Pages updates as the incident progresses.
In our example incident, the responders ask Bits to start a fix. Bits Remediation passes the relevant context to Bits Code, which creates a dedicated code channel in Slack and uses the investigation findings to generate the fix. Bits Code then creates a pull request for the responders to review.
Other incidents might call for different actions. Bits Remediation can also trigger triage actions from chat, including sending messages to teammates, paging engineers via Datadog On-Call, and creating incident or follow-up records.
Resolve the incident and capture what happened
Addressing the problem does not end the incident workflow. Responders still need to communicate the outcome, resolve the incident, preserve the investigation context, and create follow-up work. These final stages of incident response can also happen in Slack through @Datadog
.
Once the issue in our example has been addressed, Bits confirms that error rates have returned to their normal level. The responders then ask Bits to resolve the incident and create a postmortem notebook from the investigation. Because Bits already has the context from the response, the notebook documents the incident summary, relevant findings, resolution, and next steps.
The notebook gives responders a starting point for postmortem and follow-up work, removing the need to reconstruct the incident from separate conversations and tools. Creating postmortems helps teams and Bits respond to future incidents more quickly.
Start managing incidents in Slack with Bits Chat
Bits Chat in Slack brings investigation, collaboration, and action into the conversation where incident responders are already working. From the initial alert through remediation and follow-up, teams can handle incidents in Slack without having to move between tools to coordinate the response. To get started, follow the Bits Chat setup instructions for Slack. You can also learn more about using Bits Investigation to investigate issues and using Bits Code to generate code fixes.
If you don’t have a Datadog account, you can sign up for a 14-day free trial to start using Bits Chat in Slack.
Facts Only
* Nancy Zhu, Evan Marcantonio, Nicole Parisi, and Chris Miller are the listed authors.
* Bits Chat is a Datadog natural-language interface integrated into Slack.
* The tool allows users to trigger investigations using the command @Datadog investigate.
* Bits Investigation uses telemetry data, runbooks, and past incidents to form hypotheses.
* Bits Remediation suggests next steps and can trigger incident workflows or Status Page updates.
* Bits Code generates pull requests for fixes in a dedicated Slack code channel.
* Bits Chat can page engineers via Datadog On-Call and create follow-up records.
* The system generates postmortem notebooks based on the response context.
* A 14-day free trial is available for those without a Datadog account.
* Setup instructions for Slack are provided by Datadog.
Executive Summary
Datadog has integrated its Bits AI natural-language interface into Slack via Bits Chat to reduce context switching for incident responders. When a production alert triggers, the tool allows teams to initiate investigations, analyze telemetry data, and coordinate fixes within a dedicated Slack channel. This integration aims to consolidate the workflow by bringing evidence and tooling directly into the communication hub.
The process moves from automated investigation—where Bits AI forms hypotheses based on telemetry and runbooks—to remediation, which can include generating pull requests via Bits Code or updating Status Pages. Finally, the tool automates the creation of postmortem notebooks by synthesizing the incident's history. While the system streamlines the transition from alert to resolution, the effectiveness of these automated suggestions depends on the quality of the underlying telemetry and the team's willingness to review AI-generated code.
Full Take
The strongest version of this narrative is that AI-driven orchestration can solve the "swivel-chair" problem in DevOps, where cognitive load is increased by jumping between disparate tools during high-stress outages. By centering the response in Slack, the tool attempts to preserve the temporal and social context of a crisis.
However, this is a classic vendor advertorial. The narrative relies on the Authority Game by positioning Datadog’s own tool as the primary evidence for the solution to a problem it defines. It frames the "context switching" pain point as a crisis that can only be solved by deeper integration into their ecosystem. The most significant unstated assumption is that AI-generated hypotheses and code fixes are reliable enough to be initiated and reviewed within a chat interface, potentially bypassing more rigorous external validation steps.
The root paradigm is the "Single Pane of Glass" philosophy, which seeks to maximize vendor lock-in by absorbing every stage of the software development lifecycle into one platform. The second-order consequence is a potential degradation of human agency; as responders rely on Bits to "investigate," the muscle memory for manual telemetry analysis may atrophy, leaving teams vulnerable if the AI fails or hallucinates during a critical failure.
Patterns detected: ARC-0032 Authority Game
Bridge Questions:
1. How does the error rate of AI-generated "root cause findings" compare to human analysis in complex, non-linear systems?
2. What happens to the incident response process if the integration layer (Slack) or the AI orchestrator (Bits) is the cause of the production outage?
3. Does consolidating the "evidence" and the "conversation" into one stream reduce noise, or does it create a confirmation bias loop?
Counterstrike Scan: A coordinated influence campaign would manufacture a sense of "incident responder burnout" to push a specific tool as the only cure. While this content uses marketing framing, it lacks the systemic manipulation or manufactured panic typical of an attack pattern.
