I would argue that you are! You will not try to clash with reality the same way you did before, provided you “remember” and I believe future agents/models will have this kind of contextual memory continuously being getting baked in to improve..just a thought.
I think you could do this with an open model with overnight tuning on the day's errors. Probably very expensive though. Easier to scoop up all the errors on the internet on the first round of pre-training.
How do you "just" accurately evaluate the error state space of an LLM relative to a real business process? Sounds approximately impossible to me.
If you already have the business process robustly defined as code, then the utility of LLM is unclear. The value prop of LLM is in fuzzy business processes like parsing arbitrary helpdesk tickets.
You evaluate it the way we've evaluated production ML for years, with cheap QC layers sampled and checked by more expensive layers (with humans on top.)