Define what the system may prepare
Start with suggesting a category for an incoming message, for example. A suggestion must not quietly become permission to amend customer records. Identify actions outside the task and who has authority to perform them.
If structured fields or clear rules already solve the problem, they may be simpler. AI becomes a candidate where meaning or variation in the input matters. Permitted actions and access rights still need to be defined in advance.
Separate suggestion, review and execution
| Stage | Assistance | Responsibility |
|---|---|---|
| Prepare | Suggest a category and relevant passage | Keep the original and uncertainty visible. |
| Review | Show the suggestion beside source information | An authorised person checks intent and permission. |
| Execute | Carry out only the approved change | Access controls and process conditions still apply. |
| Verify | Compare the outcome with the instruction | Follow up missing or incorrect processing. |
Make human review practical
Give the reviewer the original message, relevant records and a clear explanation of why the case needs attention. An approve button is a weak control when the evidence is missing. Allow time for review and make rejecting or returning a case as practical as approving it.
Decide what happens during absence. A queued case should not become automatically approved merely because a timer expires. Keep it visibly pending or follow an agreed escalation route.
Test the mistakes that matter most
Build a test set with ordinary messages, ambiguous wording and requests outside the intended task. Define the expected handling in advance. Look beyond average scores to incorrect suggestions accepted without being noticed.
Choose any thresholds according to the task and consequences of failure. Reassess outcomes when the model, instructions or document stream changes. A threshold used in a demonstration does not establish that your process is adequately controlled.
Distinguish a score returned by a recognition service from a confidence percentage written by a language model. Do not treat a self-reported percentage as measured reliability. Use your own test cases to establish what a score means for the task; not every service provides a useful confidence score.
Make these decisions explicit
- Which evidence can a reviewer use?
- Which data may be sent to the selected service?
- Who may authorise the subsequent action?
- How is an error traced and corrected?
- When is automated assistance disabled?
Messages and documents remain input, not authority to change workflow rules.
A useful trial does not have to change records
During an initial trial, run suggestions alongside the existing process without letting them modify records independently. Review differences with the process owner. This reveals where assistance helps and where clearer instructions or a simpler approach would be preferable. Expand authority only after an explicit decision.
What would you like to clarify?
Describe the process step or open question. We can explore a practical next step together.
Discuss your process