AI / Agentic AI Interview Questions
How can you optimize a multi-agent system to reduce token cost and latency?
- Favor Plan-and-Execute over ReAct for sub-tasks with predictable steps, avoiding repeated reasoning cycles
- Use selective tool injection so each agent only sees the tools relevant to its specific role, reducing context bloat
- Cache and reuse retrieval results across agents working on the same task, instead of each agent independently re-querying the same data
- Set sensible turn budgets per agent so a stuck sub-agent fails fast rather than looping expensively
- Run independent sub-tasks concurrently rather than sequentially, when they don't depend on each other's output
- Route through a centralized gateway for tool access, reducing the redundant metadata injected by many separately connected servers
Most of these optimizations come down to the same underlying idea, give each agent only as much reasoning freedom, tool visibility, and retry budget as its specific sub-task actually needs, rather than a uniformly generous default applied everywhere.
More Related questions...