Versus Incident

Versus Incident The self-hosted AI SRE agent

UPCOMING FEATURES DEVOPS AGENT CHATOur goal is to develop an agent that can understand everything in your system, help y...
31/08/2026

UPCOMING FEATURES DEVOPS AGENT CHAT

Our goal is to develop an agent that can understand everything in your system, help you quickly catch up on the system, and make it easy to diagnose problems when they happen.

How do you know which AI model is best for incident analysis in your system? The answer is: how are you evaluating this ...
19/08/2026

How do you know which AI model is best for incident analysis in your system? The answer is: how are you evaluating this model?

The book below explains how to evaluate AI models and choose the best one for your system.

Everyone can build an agent now. Almost nobody can tell you whether theirs is getting better.You reword the prompt, swap...
15/08/2026

Everyone can build an agent now. Almost nobody can tell you whether theirs is getting better.

You reword the prompt, swap the model, add a tool. It answers — fluent, confident, plausible. Did that change help?

You can't say. Last week's incident is gone: the deploy is old, the pool is back to normal, the logs rotated. So you skim two outputs, decide it "seems better," and ship. That's tuning by guesswork, and most of us are doing it in private.

The fix isn't a smarter model. It's the test suite tool-using agents never had:

1. Record one case whose answer you already know.
2. Replay it to the agent — same data shapes, frozen scene.
3. Score what it did, not how it sounds.

Four checks do most of the work: right cause, right evidence, rejected the planted distraction, stayed inside the step budget.

The evidence one matters more than it looks. An agent that names the right cause but never opened the deploy log got lucky — and luck doesn't survive the next incident. Grading only the final answer can't tell those two apart.

I'm writing the whole thing up as a short book: a small evaluation harness in plain Python, no framework, built end to end on a DevOps incident agent and ready to point at your own.

Google published a new open spec this year, and it's almost boring on purpose.It's called the Open Knowledge Format (OKF...
13/08/2026

Google published a new open spec this year, and it's almost boring on purpose.

It's called the Open Knowledge Format (OKF) — no SDK, no database, no vendor lock-in. Just a directory of Markdown files with a small YAML header on top.

The whole idea fits in three steps:

1) One file = one concept — a YAML block on top (only `type` is required), free-form Markdown body below.

2) Many files = one bundle — just a folder. `index.md` and `log.md` are the only reserved names, everything else is a concept.

3) Files link to each other — ordinary Markdown links turn the folder into a small knowledge graph.

That's the whole format. If you can cat a file, you can read it. If you can git clone a repo, you can ship it.

What I like most: it's not a bet on any single tool. The files are already useful with nothing else running — point an AI agent at the same folder later, and it just becomes searchable by meaning too.

Breakdown below ⬇️

Compare two AI SRE tools: OpenSRE and Versus SRE Agent.An honest look at an open-source SRE agent, where it beats us, wh...
08/08/2026

Compare two AI SRE tools: OpenSRE and Versus SRE Agent.

An honest look at an open-source SRE agent, where it beats us, where we beat it, and the one idea we’re taking from it.

How can we 𝗿𝗲𝗱𝘂𝗰𝗲 𝗮𝗹𝗲𝗿𝘁 𝘀𝗽𝗮𝗺? When the same low-signal alert fires over and over, it hides the important notifications.N...
04/08/2026

How can we 𝗿𝗲𝗱𝘂𝗰𝗲 𝗮𝗹𝗲𝗿𝘁 𝘀𝗽𝗮𝗺? When the same low-signal alert fires over and over, it hides the important notifications.

New feature of Versus SRE: Alert fatigue keeps your on-call channel clean: once an alert has repeated several times in a short window, Versus treats it as “spam” and quietly sends the repeats to a separate fatigue channel instead of the channel your on-call watches. The noise is still recorded and still reachable — it just stops interrupting the people on call.

Your team already wrote the runbook to handle the incident, but when it happened, you couldn't find the fix, which is wh...
29/07/2026

Your team already wrote the runbook to handle the incident, but when it happened, you couldn't find the fix, which is why you need to teach your agent to find it the moment the incident fires. No human mistake.

The Versus SRE Agent's new feature upgrades its detection mode by integrating AI for smarter analysis, replacing the raw...
24/07/2026

The Versus SRE Agent's new feature upgrades its detection mode by integrating AI for smarter analysis, replacing the raw alerts from the SRE Agent.

We can now automatically detect real problems with the agent, allowing us to focus on developing and improving application features.

The new feature allows the SRE agent to monitor CloudWatch Metrics as a Data Source.Think of it as a teammate who watche...
19/07/2026

The new feature allows the SRE agent to monitor CloudWatch Metrics as a Data Source.

Think of it as a teammate who watches your AWS dashboards: first, they learn your service's usual rhythm, then they escalate only new or unexpected issues.

By using SRE to monitor our system, we can focus on developing application features and improving the system, instead of watching the dashboard all day.

Stop monitoring your predictions. Start monitoring the system.
18/07/2026

Stop monitoring your predictions. Start monitoring the system.

Address

Ho Chi Minh City
70000

Alerts

Be the first to know and let us send you an email when Versus Incident posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Shortcuts

Share