SyncAI.news, a Varaisys broadcasting
Context Engineering Decides Whether Your AI Testing Agents Are Useful
MB

Mayank Bhola, Forbes Councils Member

· 1 min read

World NewsForbes: Innovation

Context Engineering Decides Whether Your AI Testing Agents Are Useful

Mayank Bhola, Co-Founder and Head of Products, TestMu AI (Formerly Lambdatest).

​Every QE leader who has piloted AI test generation has some version of the same story. The agent produced an impressive volume of tests in minutes, but on closer inspection, many checked things nobody cared about, referenced UI elements that did not exist or retested the same happy path in slightly different ways.

The instinct is to blame the model. A better diagnosis is that the agent was working blind. It knows what a login page generally looks like but not what your login page does, how users depend on it or where it broke last quarter.

I have spent a decade building testing infrastructure and the last few years building testing agents, and I now believe the gap between disappointing pilots and production value has one name: context engineering. It is becoming one of the highest-leverage skills in quality engineering, and most teams have barely started practicing it.

Why Testing Agents Fail Without Context At Scale

Applied AI spent its first few years focused on prompt engineering, the craft of phrasing instructions so a model responds well. As agents began taking on multistep work, the focus shifted.

Andrej Karpathy gave the shift a clear articulation in 2025, arguing that the real work is filling the model’s context window with the right information: instructions, retrieved knowledge, memory, tool definitions and prior outputs, structured so the model can use them.

A testing agent depends heavily on knowledge that never appears in a model’s training data. Which user journeys drive revenue? Which module has a history of regressions? What did the requirement actually mean rather than what the ticket happened to say? Which failures are tolerable, and which lead to an executive escalation?

Strip that away, and the agent defaults to the statistical average of every application it has seen, producing tests for a generic app rather than yours.

Original source

This story was published by Forbes: Innovation and written by Mayank Bhola, Forbes Councils Member. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on forbes.com

Similar News