It's time to let the dog out
For decades, observability has worked like a watchdog: it barks when something's wrong and waits for you to come running. That stopped making sense the moment we could build a system that acts on its own. Why I named the company DataAgent.
Ishay YaariCo-founder & CEO
The short version
- Observability was built as a watchdog: it barks, then waits for a human to investigate and fix. That made sense only while a human had to be the one to act.
- The economics broke too. Teams pay millions a year to copy, store, and analyze logs nobody reliably reads.
- DataAgent reads live system state, topology, and drift where that data already lives, finds the real cause, and fixes it, often before anyone would have been paged.
- It stabilizes first and runs root-cause analysis second, and it only acts unsupervised on fault classes where it has already been proven right.
Why I named the company DataAgent: because it's time to replace the dog with an agent.
Why observability still works like a watchdog
For decades, observability has worked like a watchdog: it barks when something's wrong and waits for you to come running. That model made sense when a human had to be the one to investigate and fix things. It stopped making sense the moment we could build a system that acts on its own. And underneath the naming is a harder business truth: charging customers millions of dollars a year to copy, store, and analyze logs nobody reliably reads doesn't make sense anymore. Not technically, and not economically.
DataAgent is built around that realization. It's designed to do two things at once: collapse Mean Time to Resolution and slash the cost of observability itself. Instead of paging a human to stare at a dashboard and hunt for a root cause, DataAgent reads live system state, topology, and drift, finds the actual cause of an incident, and fixes it, often before anyone would have been paged. And because it reads that data where it already lives instead of re-ingesting a second, paid copy into someone else's cloud, it does this at a fraction of what teams pay today for commercial observability.
Why the incumbents can't just cut their prices
There's a real business opportunity here. Yes, it's a red ocean. But the incumbents are stuck: their revenue is the ingest, so they can't halve their annual fees without cutting their own throats.
Why an AI chat layer on a data lake isn't autonomy
The new AI-based entrants are attacking the problem from the wrong angle, bolting an AI chat interface onto the same human-shaped data lakes. Some of them do now act on their own; Datadog's Bits AI SRE went generally available in June 2026. So the question isn't whether an agent acts. It's what the agent is acting on. An LLM reasoning over pre-computed metric aggregates is deciding from evidence that was down-sampled for a human, sometimes days earlier. It can summarize what happened and suggest what to try. What it cannot do is see the fault at full fidelity at the moment it occurs.
Fix first, diagnose second, and earn autonomy one fault at a time
Coming from infrastructure automation (I was part of Cloudify, acquired by Dell in 2023), we understand what the chat-on-a-dashboard crowd doesn't: real autonomy depends on high-fidelity signals fetched on demand, exactly when a fault occurs, not a centralized lake that was down-sampled for a human days ago. It also depends on the order of operations. Most tools alert you, then leave a human to hunt for root cause before anything gets fixed. DataAgent flips that: it stabilizes the incident with a verified fix first, then runs deeper root-cause analysis in the background. No war room for problems already solved. Autonomy is earned incrementally, what we call our Trust Ladder: the system only acts on its own once it's proven right on that class of fault; every other case still routes to a human.
We also know that organizations, enterprises especially, aren't going to let a generic LLM loose on their mission-critical environments. So we're building what we call a Digital Immune System for cloud infrastructure. Think of it as a private, enterprise-grade Claude, purpose-built for SREs. Not smarter alerts. Not a better dashboard. Something that actually understands your environment, learns from it, and acts on it: autonomously, before or after things break.
The difference from everything else out there: We're not trying to filter the noise. We're eliminating it!
- #Observability
- #APM
- #Autonomous SRE
- #AIOps
- #Digital Immune System
See remediation-first in your own stack
Install an agent and watch DataAgent map your topology. No credentials, no commitment.
Keep reading
Engineering13 minLeft of Boom, Right of Boom: Prevention and Remediation Are One Loop
The industry split reliability into two disciplines: prevent it, and survive it. Both ask the same causal question — which change produces which failure — and neither can answer it alone. Why the postmortem is a broken transfer mechanism.
Nati Shalom • Aug 26, 2026
Engineering14 minThe Great Inversion: AI Writes the Code, Production Pays the Review Bill
Generation got two orders of magnitude cheaper; human verification capacity did not move. GitClear's 2026 data shows duplication at its highest recorded level while refactoring falls — and the defects that reach production are structural, not syntactic.
Nati Shalom • Aug 26, 2026
Engineering12 minPrecision vs. Recall: Why Observability Costs So Much
Observability bills are shaped by one unstated assumption: a human might need to scroll back. That is collection optimized for recall. An agent present at the moment of the fault optimizes for precision — and precision scales with your incident rate, not your estate.
Nati Shalom • Aug 26, 2026