Agents fail in production for boring reasons
The demo works because a human chose the inputs. Production is what happens when nobody does. Almost every agent we have been asked to rescue failed on one of three things, and none of them were the model.
Access. The agent needs the same permissions as the person whose job it does, which nobody wants to grant, so it gets a read-only key and quietly becomes a suggestion box. Decide the permission question first — if the answer is no, the honest scope is a drafting tool, and it should be sold internally as one.
Evaluation. If you cannot say what a good output is, you cannot tell whether last week made things better. Write twenty real cases with agreed answers before building anything, and run them on every change. It is unglamorous and it is the difference between an agent you can improve and one you can only argue about.
Ownership. An agent is a colleague who never leaves, which means someone has to be responsible for its work. Name that person on day one, give them the ability to turn it off without permission, and put its output where their existing work already lives rather than in a new dashboard nobody opens.
Get those three right and the model choice becomes a cost decision you can revisit quarterly. Get them wrong and no model will save the project.
Write the twenty test cases before the first line of code. Everything else is negotiable.
Facing this now? Send a paragraph and we will tell you where we would start.
Start an inquiry