Getting an agent to complete a scripted task in a sandbox is exciting, but hooking it up to real backend logic and production repos is an entirely different beast. You watch it handle edge cases cleanly in testing, only for orchestration to break the second it hits real API boundaries or multi-step execution.
Most of the friction isn't prompting anymore, it's managing the state and system around the agent so it doesn't drift.
Where does your setup usually fall apart when moving from a cool prototype to a reliable production workflow?