Most teams that try AI agents for operations work hit the same wall: the agent is either too unreliable to trust or too expensive to justify. This interview covers five months of running Devin, Cognition's autonomous AI engineer, on real monitoring incidents across two client projects. It walks through how the team solved reliability first by mining playbooks out of closed tickets instead of writing documentation from scratch, and only then went after cost, cutting it by 85% without giving back any of the quality gains. The order those two problems get solved turns out to be the whole story.
Background & experience:
With over 20 years of experience in IT systems and leadership, Viacheslav is currently focused on support and maintenance team structure, enterprise IT operations, and agentic AI initiatives.
Let's start with the headline numbers. Over these five months, what changed in terms of results, both quality and cost?
Viacheslav: We ran Devin for five months across two client projects, using it to investigate monitoring events the kind of alert that fires and someone has to figure out what's going on before it becomes an incident. Over that period, we deliberately reworked how the agent was given context, and two numbers moved as a result.
Investigation quality — rated 1 to 4 by an operations engineer and a project expert — went from 2.43 to 3.22. Cost per session went from $11.10 to $1.65. That's an 85% cost reduction, and it happened alongside a quality improvement, not as a trade-off.
The second number only exists because of the first. We didn't cut cost by cutting corners. We first fixed why sessions were unreliable, and once quality was solid, we could pursue cost without putting it at risk. If we'd cut cost first and quality had dropped, this wouldn't be a story worth telling. It's the sequencing that matters.
Every conversation about AI in IT operations seems to hit the same wall: "our processes aren't documented well enough." How did you get past that?
Viacheslav: We got past it by realising the objection was based on a false premise. The documentation was there. It just wasn't where people expected to look.
Everyone assumes documentation means a wiki page or a runbook someone sat down and wrote. But resolved tickets with engineer comments already hold the real diagnostic path: which dashboard, which query, what got ruled out, in what order. That's the actual trace of how the work gets done.
Why is ticket history more reliable than a document someone writes deliberately?
Viacheslav: Because a document written from scratch captures how people think the work is done. Ticket history captures how it's actually done. Those aren't always the same thing. When an engineer is resolving a real incident under time pressure, the shortcuts they take, the false leads they rule out, the specific query that finally surfaces the answer — that's the real process. A retrospective write-up tends to smooth all of that into something cleaner and less useful.
So instead of commissioning documentation, we extracted playbooks from what already existed. We started with monitors that already had runbooks, specifically because it gave us faster onboarding and a clean baseline to check the artificial intelligence's output against.
What turned out to be the hard part of building these playbooks?
Viacheslav: Domain logic. Not the mechanics of writing a playbook, but the actual substance of what needed to go into it.
No model generalises over your domain concepts on its own. For example, our system had to be told explicitly that a business event lives in a specialised traceability-history entity, not the transactions table — which is the table anyone, human or model, would instinctively reach for first. That kind of thing has to go into the playbook in plain language. The agent doesn't infer it from general reasoning.
Did the playbooks work immediately?
Viacheslav: No, and we didn't expect them to. Expect a curve, not a single pass. Each playbook improved over several sessions against real traffic. You write a version, run it against live incidents, see where it breaks down or leads the agent astray, and refine it.
Six weeks of that rhythm moved the share of sessions meeting or exceeding our human baseline from 53% to 81% on one project, and from 42.5% to 93.8% on the other. That second number in particular tells you the ceiling on this approach is high once the domain logic is actually captured.
So quality was solved. Then what went wrong?
Viacheslav: Then we looked at the bill. At 3.7 ACUs per session, the cost annualised to roughly $4,600 and $3,100 per project. That's more than what it replaced. The agent was faster, more consistent, and available at 3 a.m., genuinely better on every quality dimension, and we were still losing the business case on economics alone.
What did you try first to bring that cost down?
Viacheslav: Everything within our control. Tighter playbooks, cleaner configuration, more discipline about which tools the agent was even allowed to reach for. That work got us from 3.7 ACUs down to 2.7 ACUs.
That's a real improvement of about 27%, but it was still the wrong order of magnitude. We needed something structural, not incremental.
That's when you went to Cognition directly?
Viacheslav: Yes, and plainly: we told them we couldn't run their product at this price. The first request went nowhere.
That's the point where most teams quietly conclude AI is too expensive and stop pursuing it. We didn't stop there. We rewrote the case, pushed harder, escalated it, and eventually got real engineering help back from their side.
Why keep pushing instead of accepting the pricing as fixed?
Viacheslav: Unit economics is a legitimate support case, not a pricing fact you're supposed to accept in silence. Vendors don't always have full visibility into how a workload actually behaves in production, or what's technically possible on their own stack, until a customer forces the conversation. Treating the sticker price as non-negotiable data, rather than a starting point for a technical discussion, is how teams talk themselves out of viable tools.
What came out of that escalation?
Viacheslav: Three recommendations. Two didn't apply to our situation. The third did: SWE-1.7, Cognition's experimental model, which is now available as the "lite" option in Devin's settings.
Switching to it dropped consumption to 0.55 ACUs per session. Annual cost fell to $686 and $458 for the two projects. And quality held — it didn't regress.
That's the part that seems counterintuitive. Why didn't quality drop when you moved to a smaller, cheaper model?
Viacheslav: Because it wasn't luck, and it's not a claim that small models generally beat large ones. It's specific to what had already happened by that point.
Powerful models earn their price by compensating for weak context. When an agent has to reason its way through ambiguity — figure out which table to check, what to rule out, what the domain even means — you need a model with more raw reasoning capacity to get there.
By the time we made the switch, our playbooks had absorbed weeks of iteration. We'd already written down the domain logic explicitly. The agent had very little left to work out on its own. We'd effectively been paying frontier-model rates for reasoning that the playbooks had already made unnecessary.
If you had to leave someone with one operating principle from this, what would it be?
Viacheslav: Right-size the model to your context, not to how hard the task looks on paper.
A task can look hard in the abstract and still be cheap to run well, if the context handed to the model is strong enough. The cost of a given unit of work keeps falling on its own. Cheaper models get more capable every cycle, and that trend isn't something you have to work for.
Captured domain knowledge doesn't get cheaper by itself, though. Someone has to extract it, write it down, and iterate on it against real cases. That's the part that decides how much of the cost curve you actually get to capture. It's also the part every team already has sitting in their closed tickets, whether they've looked there or not.
FAQs
An AI agent can take multi-step actions toward a goal, investigating an issue, querying systems, ruling out causes rather than just answering questions in a single turn. In IT operations, this means an agent like Devin can work through a monitoring alert the way an engineer would: checking dashboards, running queries, and narrowing down a root cause, instead of just summarising information a person hands it.
It depends a lot on how much reasoning the agent needs to do compared to how much context it gets up front. Cost is usually measured per session or per compute unit (ACUs, for Devin). It can also change a lot based on the model you use and how well your processes are documented. Teams that put domain knowledge into playbooks first can often switch to smaller, cheaper models later without losing quality. This makes a costly pilot sustainable.
Related insights
The breadth of knowledge and understanding that ELEKS has within its walls allows us to leverage that expertise to make superior deliverables for our customers. When you work with ELEKS, you are working with the top 1% of the aptitude and engineering excellence of the whole country.
Right from the start, we really liked ELEKS’ commitment and engagement. They came to us with their best people to try to understand our context, our business idea, and developed the first prototype with us. They were very professional and very customer oriented. I think, without ELEKS it probably would not have been possible to have such a successful product in such a short period of time.
ELEKS has been involved in the development of a number of our consumer-facing websites and mobile applications that allow our customers to easily track their shipments, get the information they need as well as stay in touch with us. We’ve appreciated the level of ELEKS’ expertise, responsiveness and attention to details.