
In short
After six months using specialised AI agents for machine learning, backend, frontend, DevOps, and QA work, an engineering team reports faster delivery and debugging. It manages agent limitations by assigning clear roles, modularising the codebase, versioning prompts, restarting sessions, and dividing large tasks into focused pieces.
Six months ago, we handed HoverBot's repo to a squad of AI agents. Today, they feel like full-time teammates.
These agents help us ship faster, refactor with confidence, and rethink traditional workflows. We treat them like a real engineering team, each with a clear role:
- ML engineer agent fine-tunes models and guards data pipelines
- Backend agent ships APIs and core logic
- Frontend agent shapes UI and application state
- DevOps agent owns Docker, CI/CD, and infra
- QA agent generates and runs automated tests
This setup helps us scale more effectively, debug faster, and keep responsibilities clean and clear.
Of course, it's not without challenges:
1. The ecosystem evolves fast.
AI coding agents update every week. Models improve quickly as well, but sometimes become less predictable. What worked last week might break today. So we treat prompts and tools like versioned APIs: something we regularly test, review, and maintain.
2. Context is still a big limitation.
Even with models going from 4K to 200K tokens, when you work with large codebases, things get messy. Important details get lost. Outputs can get fuzzy. We've found the best way to handle this is by restarting sessions often and keeping tasks small and focused.
3. Testing works differently.
Traditional unit tests don't translate well to agent workflows. As engineers, we're used to writing lots of small tests that cover each logical branch. This makes refactoring safe and keeps things reliable, even if it means our test codebase is bigger than the implementation. But with AI agents, every extra test consumes context. So instead of classic unit tests, we've started using retrieval-based test scaffolds and lightweight validation checks after generation.
What's Working Well
Here's what's been working well for us:
- Keeping the codebase modular and minimal
- Defining clear task boundaries for each agent
- Versioning prompts (PromptOps is real!)
- Breaking big tasks into smaller, manageable chunks
Right now, the tools are starting to catch up. Our development speed has improved significantly. And multi-agent workflows aren't just an experimental idea anymore, they're becoming part of how we actually build software.
Frequently asked questions
- What roles can AI agents perform on a software development team?
- The team assigns separate agents to machine learning, backend, frontend, DevOps, and quality assurance work. Their responsibilities include fine-tuning models, guarding data pipelines, shipping APIs and core logic, shaping interfaces and application state, managing Docker, CI/CD, and infrastructure, and generating and running automated tests.
- How can developers manage AI agent context limits in large codebases?
- Large codebases can cause important details to be lost and outputs to become less precise, even with larger context windows. The team addresses this limitation by restarting sessions frequently, maintaining a modular and minimal codebase, defining narrow task boundaries, and dividing large assignments into smaller, focused pieces.
- How does testing change when AI agents generate code?
- Every additional test consumes context available to an AI agent. Instead of depending entirely on numerous small unit tests covering individual logical branches, the team uses retrieval-based test scaffolds and lightweight validation checks after generation. This approach reflects the different context requirements of agent-driven development while retaining checks on generated output.
- What practices help multi-agent software development work effectively?
- The workflow keeps the codebase modular and minimal, gives every agent a clear responsibility, versions prompts, and breaks large tasks into manageable chunks. Prompts and tools are treated like versioned APIs that require regular testing, review, and maintenance because coding agents and their underlying models can change quickly.
About the author
Founder & CEO at HoverBot
Founder of HoverBot, where he leads product strategy and applied AI architecture, and CTO and co-founder of WTFox.ai. Nineteen years in software engineering, most recently as Software Architect at Mercer, where he shipped HR chatbots and OCR claims processing on Azure AI, and as tech lead at Darwin and Technosoft SEA, after engineering roles at Sberbank, Veon, and Softline. Hands-on with architecture decisions, deployment operations, and benchmark-driven quality optimization. Based in Singapore.
- 19 years of software engineering, architecture, and engineering leadership
- Founder of two AI startups: HoverBot and WTFox.ai
- Applied AI: conversational systems, RAG pipelines, agentic workflows, and safety controls
- Enterprise AI delivery: HR chatbots and OCR claims processing on Azure AI at Mercer
- Led engineering teams of 10+ as tech lead and software architect
- Cross-industry: enterprise HR and benefits, banking, telecom, automotive, e-commerce and marketplaces
- Writes on AI chatbot architecture, agentic systems, and deployment patterns at vitaliks.me


