Begin by clarifying the user outcome, workload and failure costs. State the assumptions you need: request volume, latency target, data freshness, access model and whether the system may abstain. Do not invent precise scale numbers when the prompt has not supplied them.
Draw the simplest complete request path. For an AI application, this may include ingestion, retrieval, generation, tools and delivery. Identify where state lives and which transitions must be durable. Add the model only where its behavior serves a clear requirement.
Pick two or three hard decisions and explain the alternatives. A candidate-generation stage trades recall against cost; a cache trades freshness against latency; an asynchronous workflow changes the user interaction and failure recovery. Explain the constraint that selects your choice.
Finish with verification and operations. Name the evaluation dataset, error categories, telemetry and rollback condition. Be explicit about what an offline test cannot prove. If time is short, use one concrete failure scenario to demonstrate that the components form a coherent system.
Practice aloud: spend one minute defining requirements, three minutes on the request path, three minutes on tradeoffs and two minutes on failure handling. Review whether every claimed guarantee has a mechanism behind it.