PARAPHRASED · ANONYMISED · 77 QUESTIONS
Real interview questions
Technical questions from real engineering hiring conversations, mostly for AI roles, paraphrased and anonymised. Each shows the type of company, the role and the year – never the company’s name.
How these were collected
Each question is paraphrased from a real technical hiring conversation, mostly for AI and ML engineering roles. Names of companies, people and products are removed; each question keeps only the type of company, the role and the year.
We limited how many questions come from any one company, and there is no way to filter by company.
“Asked in N interviews” means the same question came up in more than one interview; the label names one of the companies.
Showing 77 of 77 questions
RAG & retrieval
7 questionsHow did you chunk documents for your RAG pipeline, and what drove that chunking strategy?
Asked at a travel tech startup · Full-Stack Engineer · 2026
Users report that your LLM assistant confidently returns wrong numbers when answering questions about business metrics. How do you work out where the errors come from, and how do you fix them?
Asked at an e-commerce scale-up · AI Engineer · 2026
An LLM answers questions over data stored in Postgres. When would you add a graph database or model knowledge as subject-predicate-object triples, and what data and query patterns make that genuinely better than relational tables?
Asked in 7 interviews, including at a proptech startup · Founding Engineer · 2026
You say you cut retrieval latency on a production RAG system. Walk through how you found the cause. Was the bottleneck in chunking and ingestion or in retrieval itself, and why did your change actually improve retrieval performance?
Asked in 3 interviews, including at a legaltech startup · Applied AI Engineer · 2026
Why still embed documents when an agent can grep and search the corpus directly? For about 10,000 legal documents, would you choose lexical search, embeddings, or both, and why?
Asked in 2 interviews, including at a legaltech startup · Applied AI Engineer · 2026
When is vector RAG a poor fit? With fine-grained document permissions, pre-filtering can slow HNSW search while post-filtering can prune results to nothing. How would you enforce access control without wrecking recall or latency?
Asked in 2 interviews, including at a legaltech startup · Applied AI Engineer · 2026
For an LLM system that explains why genes are linked to disease mechanisms, would you add a knowledge graph of biological pathways on top of RAG? What does the graph buy you, and what expertise would the build need?
Asked in 2 interviews, including at a biotech startup · Contract AI Engineer · 2026
Agents & tool use
19 questionsWalk me through how you build software with AI today, from planning to pull request. Which coding agents, CLIs or MCP servers do you rely on, and how do you make sure generated code is correct and polished enough to ship?
Asked in 5 interviews, including at an enterprise AI software startup · Product Engineer · 2026
What makes it hard to take an agentic AI system from proof of concept to full-scale production? Is it just a longer timeline, or does the engineering work itself change, and how?
Asked at a marketing & adtech startup · AI Engineer · 2026
For an internal workflow automation, how do you decide whether you need an LLM at all? When would you use a workflow engine that calls an LLM only for the semantic-matching steps, rather than an LLM end to end?
Asked in 2 interviews, including at an established digital media company · AI Automation Engineer · 2026
Before hiring for a role such as sales, a founder asks whether an AI agent could do part or all of the job. How would you assess whether an agent can match a human hire, and whether it is worth building?
Asked in 2 interviews, including at a consumer tech startup · Founding Engineer · 2026
Coding agents usually give the model a general shell tool, yet they also ship dedicated read-file and write/edit-file tools. Why provide those when the shell can already do the same thing? What are the benefits?
Asked at an enterprise AI software startup · Software Engineer · 2026
How would you implement memory for an AI agent, and how would it make the overall system more useful?
Asked at a defence & security startup · Senior Founding Engineer · 2026
When would you choose a single-agent design over a multi-agent system? Define what you mean by each, then walk through the orchestration patterns you've used in production and the trade-offs of each.
Asked in 2 interviews, including at an HR tech & recruiting startup · Product Engineer · 2026
Should the agents you build run fully autonomously, or be designed for human-agent collaboration? How do you decide, and how would you structure the handoffs between agent and human?
Asked at an enterprise AI software company · Applied AI Engineer · 2026
What's your take on the current wave of open-source autonomous agent harnesses, and what would they need to become enterprise-ready?
Asked in 2 interviews, including at a fintech startup · AI Engineer · 2026
For an LLM that answers questions from a graph database, walk through one concrete user question end to end: the LLM's tool call, the exact graph query it runs, and why that beats modelling the same entities in Postgres.
Asked in 5 interviews, including at a proptech startup · Founding Engineer · 2026
A team already chains existing bioinformatics tools and pretrained models for its use case. What would an agentic infrastructure layer add beyond that, how does it differ from using an AI coding assistant to write pipelines, and what's the value?
Asked in 3 interviews, including at a biotech startup · Senior Computational Biologist · 2026
Tell me about structuring messy, unstructured operational data for an AI agent. How did you make it useful to an engineer querying it to diagnose what went wrong, why it happened, and what action to take?
Asked at a fintech startup · Founding Engineer · 2026
An agent must use a large, growing catalogue of third-party integrations, too many to load every tool schema into context. How does it pick the right ones and learn their inputs and outputs, and when would you use MCP rather than direct API calls?
Asked in 4 interviews, including at an enterprise AI software startup · Software Engineer · 2026
We're building many agents across different business functions. Design a core agent platform that lets teams ship new agents quickly and reliably. What belongs in the shared platform, and is centralizing these capabilities even the right abstraction?
Asked in 2 interviews, including at an e-commerce scale-up · AI Engineer · 2026
Design an AI agent that colleagues message in a workplace chat tool's channels and DMs. It acts through many third-party integrations but must never surface information from channels the requesting user can't access. Walk through the agent architecture.
Asked at an enterprise AI software startup · Software Engineer · 2026
Once the agent knows which API endpoints it needs, how should it execute them: one tool per endpoint, a few generic request tools, or writing and running code in a sandbox? Which tools would you expose, and what are the trade-offs?
Asked in 4 interviews, including at an enterprise AI software startup · Software Engineer · 2026
Design an agentic platform that gathers relevant information from across the internet using many different APIs and tools, handling varied data formats. How would you implement it, and what challenges would you expect to face?
Asked in 2 interviews, including at a startup · Founding Engineer · 2026
An agent has run for an hour. Tool outputs are already trimmed, but its context window is filling up and the task isn't done. Without adding a time limit, walk through how you keep it running instead of erroring out.
Asked in 4 interviews, including at a legaltech startup · Applied AI Engineer · 2026
What's your system design philosophy for building agentic workflows from scratch in high-risk domains like finance or insurance? How do you structure guidelines and guardrails so agents don't drift in the wrong direction?
Asked at a fintech startup · Agentic AI Engineer · 2026
LLM evaluation
9 questionsWhich LLMOps tools have you used to trace, evaluate and monitor LLM agents in production, and what did each one actually tell you about how your agents were behaving?
Asked in 2 interviews, including at an e-commerce scale-up · AI Engineer · 2026
What does a good evaluation setup for an LLM application look like in practice? What do you measure, and how do eval results change what you build or ship?
Asked at a legaltech startup · Applied AI Engineer · 2026
For an LLM system you shipped, how did you build the evals and the golden dataset behind them: how big, how collected and labelled (predefined metrics, human labels or both), why was it hard, and what gaps made you invest in it?
Asked in 5 interviews, including at an e-commerce scale-up · AI Engineer · 2026
Give an example where you had to evaluate a model against messy data with no real ground truth. How did you approach it, and how did you convince yourself the evaluation was meaningful?
Asked at a market research startup · Research Engineer · 2026
When an LLM-based system produces numerical figures, how do you check they are actually correct? Do you define a ground-truth set, and how would you build it?
Asked at a defence & security startup · Senior Founding Engineer · 2026
You're iterating quickly on an agentic AI product. How do you evaluate each new version, know it's actually better than the last, and catch regressions in existing behaviour before it ships?
Asked in 2 interviews, including at a defence & security startup · Founding AI Engineer · 2026
Your LLM provider deprecates the model version behind a production workflow, and the replacement behaves slightly differently. How do you monitor that the system still meets its KPIs, and how do you catch and manage regressions during the upgrade?
Asked in 3 interviews, including at an established digital media company · AI Automation Engineer · 2026
How would you judge an LLM agent's real-world performance when public benchmarks may mostly reward recall of training data? How can you tell whether a task is genuinely outside the model's training distribution?
Asked in 2 interviews, including at a deeptech startup · ML Research Engineer · 2025
How would you measure the uncertainty of a foundation model's output? Compare black-box approaches, such as perturbing the prompt and measuring output variance, with white-box approaches that use open-weight model internals, such as token probabilities or hidden states.
Asked in 3 interviews, including at a market research startup · Research Engineer · 2026
Prompting & structured output
1 questionYou need to order a hotel's 40–80 photos for its gallery. Can a vision-language model actually rank that many images in one pass, or is it only reliable at labeling and captioning? How would you design it?
Asked in 2 interviews, including at a travel tech startup · Product Engineer · 2026
LLM serving & cost
2 questionsAn LLM-backed processing pipeline is now in production. What would your monitoring and observability stack look like: what would you log, which health metrics would you track and alert on, and what is specific to the LLM calls?
Asked in 3 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026
Your pipeline depends on hosted vision-language model APIs. How would you handle rate limits and provider outages when 1,000 jobs arrive at once and must be processed in near real time, so queue back-pressure isn't acceptable?
Asked in 2 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026
AI safety & security
7 questionsAn LLM classifier builds its user prompt as 'Classify this document:' followed by the raw document text. What security problem does this create, and how would you mitigate it?
Asked at an AI scale-up · Applied AI Engineer · 2026
A team worries that sending proprietary research data to a third-party hosted LLM could leak it. How could exposure actually happen with a hosted model versus a self-hosted one, and how would you reduce the risk?
Asked at a biotech startup · Contract AI Engineer · 2026
In AI security, where do you draw the boundary between reliability and safety issues like hallucinations and true security issues like data protection and adversarial attacks? How would you prioritise them for a healthcare deployment?
Asked at a cybersecurity startup · Founding AI Engineer · 2026
What risks do you see in letting an AI system produce financial reporting and billing, for example connected directly to an ad platform? What guardrails would you add, and how would you validate its outputs before trusting it?
Asked at an established digital media company · AI Automation Engineer · 2026
One generalist agent serves users from different teams, each entitled to different datasets. Would you split it into specialist agents? How do you enforce per-user data authorization when the agent calls tools, and where does that layer live?
Asked in 5 interviews, including at an e-commerce scale-up · AI Engineer · 2026
If you deploy agents that act inside your customers' own enterprise systems, how would you design for governance and auditability in production?
Asked at an enterprise AI software company · Applied AI Engineer · 2026
When securing LLM agents in an enterprise, where do you draw the line between what can be formally verified and what can't? How would you specify a security or context-handling property precisely enough to verify it formally?
Asked at a cybersecurity startup · Founding AI Engineer · 2026
ML fundamentals
3 questionsIf every new evaluation example requires an expensive real-world experiment, how do you decide how many data points you need to judge a model's performance with statistical confidence?
Asked at a deeptech startup · ML Research Engineer · 2025
If a vision model scores hotel photos, will it generalize across very different properties, such as a small boutique hotel with few amenities versus a large resort with hundreds of photos? How would you evaluate and handle that?
Asked at a travel tech startup · Product Engineer · 2026
You propose adding a Bayesian prior to a model. How does that actually work, and why would it beat plain fine-tuning, which could learn the same structure directly in its weights?
Asked at a startup · Research Engineer · 2026
Model training & fine-tuning
1 questionYour system needs to understand images. Do you fine-tune current open-weight models to add that capability now, or wait for stronger open-weight multimodal models? How do you weigh the tradeoff and make the call?
Asked at a market research startup · Research Engineer · 2026
Data engineering
1 questionYou're building an AI root-cause-analysis tool for equipment failures across large facilities. Where does the data come from (maintenance logs, sensor/IoT telemetry), how do you handle its quality problems, and where does machine learning fit?
Asked in 5 interviews, including at a defence & security startup · Founding AI Engineer · 2026
ML system design
6 questionsFor an LLM product already in production, what patterns would you use to cut its latency and to build a tight user-feedback loop that actually drives model and prompt improvements?
Asked at a healthtech startup · AI Engineer · 2026
Six months after launch, your AI agent product has many customers in production and ships regular releases. Then a problem hits production. How do you work out what's going on, get things running again, and approach supporting an AI system in production?
Asked at a defence & security startup · Founding AI Engineer · 2026
Which cross-cutting concerns would you pull into shared components across agents (cost tracking, tracing and quality monitoring, chat-channel integrations, model and effort routing, workflow orchestration), and which should stay agent-specific? Why?
Asked in 3 interviews, including at an e-commerce scale-up · AI Engineer · 2026
Screen recordings and timestamped click and keystroke logs land in blob storage. Design the services that turn them into a process graph shown in a UI: the processing stages, where a VLM or LLM fits, what you'd store, and which databases.
Asked in 2 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026
How would you represent a text-heavy scientific procedure, such as a step-by-step lab protocol, so it can be embedded, compared and used efficiently by an ML model?
Asked in 2 interviews, including at a deeptech startup · ML Research Engineer · 2025
You can afford only a few dozen physical experiments. How would you design a closed loop that starts from prior results, including known failures, and uses a model to pick each next experiment to get the best outcome on that budget?
Asked in 2 interviews, including at a deeptech startup · ML Research Engineer · 2025
Software system design
5 questionsExplain the difference between client-side and server-side rendering. What are the trade-offs, and when would you choose each?
Asked at a foundation-model scale-up · Applied AI Engineer · 2026
An LLM-backed processing pipeline handles about 100 items a day. If volume grows to 1,000 a day, what exactly breaks first, which bottlenecks would you look for, and how would you adapt the design?
Asked in 2 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026
When you choose tools for an internal AI automation, why should the stack your engineering team already uses affect the decision? Which dependencies would you weigh?
Asked at an established digital media company · AI Automation Engineer · 2026
An early-stage product is built on Next.js, Supabase and GitHub. As founding engineer, would you keep that stack for what comes next or change it? What would drive your decision?
Asked at a consumer tech startup · Founding Engineer · 2026
A no-code AI agent builder covers most of a workflow, but the last 20% needs custom logic, such as bespoke filtering on a third-party data API. How do you let non-technical users build that without code and without cluttering every node with options?
Asked in 4 interviews, including at an enterprise AI software startup · Product Engineering Lead · 2023
Backend & APIs
2 questionsYour product has to integrate with customers' existing systems, often decades-old on-prem software or third-party SaaS with poor APIs. How do you plug into them reliably?
Asked at a travel tech startup · Full-Stack Engineer · 2026
You need to turn a batch LLM classification script into a scalable FastAPI service. Sketch the design: the main classes and how they interact, and the overall service architecture.
Asked in 2 interviews, including at an AI scale-up · Applied AI Engineer · 2026
Python & coding
1 questionA junior colleague's quick Python script loads documents from a JSON file, classifies each with an LLM API, and writes a CSV report. Without writing code, review it: list the issues, rank them by severity, and say which you'd fix first.
Asked in 5 interviews, including at an AI scale-up · Applied AI Engineer · 2026
Databases & SQL
1 questionUser management and CRUD data live in Postgres, while workflow graphs live in a graph database. How would you relate the two? Would you replicate or sync data between them, and how would you keep them consistent?
Asked at an enterprise AI software company · Applied AI Engineer · 2026
Cloud & infrastructure
1 questionHow would you take a system with a FastAPI backend, a React client and background processing workers to production? Would it all be one service or several, and what drives that choice?
Asked at an enterprise AI software company · Applied AI Engineer · 2026
Product sense
5 questionsYou join as an internal AI engineer and are asked to help the marketing team. How do you work with stakeholders to gather requirements and find the automation opportunities worth building?
Asked at an established digital media company · AI Automation Engineer · 2026
Are good AI automation targets always hard, hard-to-spot problems, or also routine work nobody should do by hand? How would you weigh a complex task done rarely against a simple task done constantly?
Asked in 2 interviews, including at an established digital media company · AI Automation Engineer · 2026
Imagine it's two years from now and we're discussing what surprised everyone in AI. What's the most counterintuitive development you expect, one that most people are missing today?
Asked at an HR tech & recruiting startup · Product Engineer · 2026
Generative AI could clearly improve customer-facing enterprise workflows like voice support. Why aren't more large enterprises adopting it? Who inside the organization is usually saying no, and why?
Asked at an enterprise AI software startup · Product Engineering Lead · 2023
Where do you think SaaS is headed? If AI-generated, dynamic front ends become the norm, what's your thesis on which parts of today's software and product stack survive?
Asked in 2 interviews, including at a construction tech startup · Founding AI Engineer · 2026
Technical leadership
6 questionsIndustry research can't examine every angle the way academia can. On a project with a product deadline, how do you decide the evidence is good enough to ship an MVP rather than keep investigating?
Asked at a market research startup · Research Engineer · 2026
You've shipped ten internal AI automations for non-technical business teams. Who owns and maintains them, what go-live and support model keeps them all running, and how does aiming for team self-sufficiency change the tools you'd build with?
Asked in 4 interviews, including at an established digital media company · AI Automation Engineer · 2026
You're brought in to add an LLM capability to an existing research platform. How would you define the architecture, ship a fast first iteration in sprint one, decide what it must validate, and plan sprint two?
Asked at a biotech startup · Contract AI Engineer · 2026
Suppose you join to build ML experimentation pipelines from scratch. What infrastructure, tooling and investment would you ask for up front so you can move very quickly?
Asked at a healthtech startup · ML Engineer · 2026
With a fixed early-stage hiring budget, how many engineers would you hire for the first team, and would you include juniors or only seniors? Walk through how you'd reason about the trade-offs.
Asked in 2 interviews, including at a defence & security startup · Founding CTO · 2026
Users of an AI workflow-automation platform can request custom extensions to its third-party integrations. If it grew to millions of users, that could mean thousands of requests. How would you prioritize them, allocate integration engineers, and scale the integrations team as usage grows?
Asked in 3 interviews, including at an enterprise AI software startup · Product Engineering Lead · 2023