Contents
An AI feature is part of an application with users, data, permissions and failure states. Production readiness means evaluating that whole service in its intended context. Define what the feature may do, how its results will be assessed and who responds when the system cannot complete the task acceptably.
Define the task and the consequences of error
Describe the user's goal and the role of the AI component. Is it drafting something for review, classifying an item or proposing an action in another system? State which decisions remain with a person and which actions the application is permitted to perform.
Identify the consequences of an incorrect result and involve the owners responsible for those consequences. NIST's AI Risk Management Framework is a voluntary resource for incorporating trustworthiness considerations into AI design, development, use and evaluation. It provides a wider risk-management reference; using it does not by itself certify an application as ready.
Build an evaluation around representative tasks
Assemble authorised examples that reflect the intended use, including difficult and incomplete inputs. Define what an acceptable result looks like and which failures are unacceptable. Keep the examples and review criteria under version control or another appropriate controlled process so changes can be compared consistently.
Assess the full journey as well as the model output. Check whether the application retrieves the right information, preserves the user's permissions and presents a result the user can review. Record the conditions and limitations of the evaluation rather than presenting one aggregate score as complete evidence of readiness.
Enforce permissions outside the model
Make the application responsible for deciding which records and actions a user may access. Validate tool arguments and enforce authorisation close to the action. Treat generated content and retrieved material as untrusted input when they reach another system boundary.
Limit tools to the operations required by the task and define the circumstances in which a person must approve an action. Exercise attempts to access unavailable data or perform an out-of-scope operation. A model instruction describing the boundary is not a substitute for application checks that enforce it.
Design the unsuccessful paths
Decide what happens when the provider is unavailable, a response is unusable or the required information cannot be found. Give the user a truthful status and a useful next step. Avoid presenting a guessed result as successful completion simply to keep the interface moving.
Define limits on time, retries and external work. If an operation changes data, ensure a retry cannot unintentionally repeat the action. Explain which failures can be retried and which need review, and preserve enough safe diagnostic information for the operating team to investigate.
Review changes to the whole AI feature
Track the configuration that affects behaviour, including prompts, model selection, retrieval settings and tool definitions. Run the relevant evaluation when those parts change. Separate a provider's claimed capability from evidence that the particular application behaves acceptably.
Agree release criteria and a recovery plan. Consider how to disable the feature or return to a known configuration if the result is unacceptable. Identify which parts of the surrounding service should remain usable while the AI capability is unavailable.
Own performance, cost and review after launch
Monitor the indicators needed to operate the feature, such as failed tasks, response times and resource usage. Keep personal or confidential content out of routine analytics and logs unless there is a defined, authorised reason and appropriate handling. Decide who reviews reported problems and updates the evaluation examples.
Compare real use with the assumptions made before launch. Review whether the human oversight process is workable and whether users understand the feature's limits. Treat production readiness as an operating responsibility that continues as the product and its dependencies change.
A release exercise for an AI request assistant
Consider a fictional assistant that reads a service request, drafts a category and suggests a reply. Its supported release stops at a staff-reviewed suggestion; it cannot approve expenditure or send a reply independently. Evaluate complete cases: an ordinary request, contradictory instructions, missing evidence, an inaccessible attachment and an instruction embedded in a document that asks the system to expand its permissions.
For each case record expected behaviour, observed result, reviewer correction and downstream state. “No unsupported factual statement” and “no message sent before authorised approval” are separate conditions. A well-written answer that causes an unauthorised send fails the second. Agree task-specific acceptance with the service owner rather than borrowing a universal accuracy threshold.
Before launch, exercise provider failure, stale source material, repeated approval and feature withdrawal. The existing manual journey should remain understandable when generation is unavailable. Attach the evaluation set and operating checks to the exact model, prompt, retrieval settings and application version. NIST’s AI risk framework is a voluntary risk-management reference; citing it does not certify a release or establish the feature’s fitness for use.
How this relates to Veda Software’s work
Powerleague provides adjacent evidence about an order-to-fulfilment workflow and human operating roles. It is not an AI assistant case study or proof of model evaluation. Veda Software’s role here is application engineering; VEDA AI retains strategy and adoption work.