PARAPHRASED · ANONYMISED · 77 QUESTIONS

Real interview questions

Technical questions from real engineering hiring conversations, mostly for AI roles, paraphrased and anonymised. Each shows the type of company, the role and the year – never the company’s name.

How these were collected

Each question is paraphrased from a real technical hiring conversation, mostly for AI and ML engineering roles. Names of companies, people and products are removed; each question keeps only the type of company, the role and the year.

We limited how many questions come from any one company, and there is no way to filter by company.

“Asked in N interviews” means the same question came up in more than one interview; the label names one of the companies.

Topic

Showing 77 of 77 questions

RAG & retrieval

7 questions
  1. How did you chunk documents for your RAG pipeline, and what drove that chunking strategy?

    Asked at a travel tech startup · Full-Stack Engineer · 2026

  2. Users report that your LLM assistant confidently returns wrong numbers when answering questions about business metrics. How do you work out where the errors come from, and how do you fix them?

    Asked at an e-commerce scale-up · AI Engineer · 2026

  3. An LLM answers questions over data stored in Postgres. When would you add a graph database or model knowledge as subject-predicate-object triples, and what data and query patterns make that genuinely better than relational tables?

    Asked in 7 interviews, including at a proptech startup · Founding Engineer · 2026

  4. You say you cut retrieval latency on a production RAG system. Walk through how you found the cause. Was the bottleneck in chunking and ingestion or in retrieval itself, and why did your change actually improve retrieval performance?

    Asked in 3 interviews, including at a legaltech startup · Applied AI Engineer · 2026

  5. Why still embed documents when an agent can grep and search the corpus directly? For about 10,000 legal documents, would you choose lexical search, embeddings, or both, and why?

    Asked in 2 interviews, including at a legaltech startup · Applied AI Engineer · 2026

  6. When is vector RAG a poor fit? With fine-grained document permissions, pre-filtering can slow HNSW search while post-filtering can prune results to nothing. How would you enforce access control without wrecking recall or latency?

    Asked in 2 interviews, including at a legaltech startup · Applied AI Engineer · 2026

  7. For an LLM system that explains why genes are linked to disease mechanisms, would you add a knowledge graph of biological pathways on top of RAG? What does the graph buy you, and what expertise would the build need?

    Asked in 2 interviews, including at a biotech startup · Contract AI Engineer · 2026

Agents & tool use

19 questions
  1. Walk me through how you build software with AI today, from planning to pull request. Which coding agents, CLIs or MCP servers do you rely on, and how do you make sure generated code is correct and polished enough to ship?

    Asked in 5 interviews, including at an enterprise AI software startup · Product Engineer · 2026

  2. What makes it hard to take an agentic AI system from proof of concept to full-scale production? Is it just a longer timeline, or does the engineering work itself change, and how?

    Asked at a marketing & adtech startup · AI Engineer · 2026

  3. For an internal workflow automation, how do you decide whether you need an LLM at all? When would you use a workflow engine that calls an LLM only for the semantic-matching steps, rather than an LLM end to end?

    Asked in 2 interviews, including at an established digital media company · AI Automation Engineer · 2026

  4. Before hiring for a role such as sales, a founder asks whether an AI agent could do part or all of the job. How would you assess whether an agent can match a human hire, and whether it is worth building?

    Asked in 2 interviews, including at a consumer tech startup · Founding Engineer · 2026

  5. Coding agents usually give the model a general shell tool, yet they also ship dedicated read-file and write/edit-file tools. Why provide those when the shell can already do the same thing? What are the benefits?

    Asked at an enterprise AI software startup · Software Engineer · 2026

  6. How would you implement memory for an AI agent, and how would it make the overall system more useful?

    Asked at a defence & security startup · Senior Founding Engineer · 2026

  7. When would you choose a single-agent design over a multi-agent system? Define what you mean by each, then walk through the orchestration patterns you've used in production and the trade-offs of each.

    Asked in 2 interviews, including at an HR tech & recruiting startup · Product Engineer · 2026

  8. Should the agents you build run fully autonomously, or be designed for human-agent collaboration? How do you decide, and how would you structure the handoffs between agent and human?

    Asked at an enterprise AI software company · Applied AI Engineer · 2026

  9. What's your take on the current wave of open-source autonomous agent harnesses, and what would they need to become enterprise-ready?

    Asked in 2 interviews, including at a fintech startup · AI Engineer · 2026

  10. For an LLM that answers questions from a graph database, walk through one concrete user question end to end: the LLM's tool call, the exact graph query it runs, and why that beats modelling the same entities in Postgres.

    Asked in 5 interviews, including at a proptech startup · Founding Engineer · 2026

  11. A team already chains existing bioinformatics tools and pretrained models for its use case. What would an agentic infrastructure layer add beyond that, how does it differ from using an AI coding assistant to write pipelines, and what's the value?

    Asked in 3 interviews, including at a biotech startup · Senior Computational Biologist · 2026

  12. Tell me about structuring messy, unstructured operational data for an AI agent. How did you make it useful to an engineer querying it to diagnose what went wrong, why it happened, and what action to take?

    Asked at a fintech startup · Founding Engineer · 2026

  13. An agent must use a large, growing catalogue of third-party integrations, too many to load every tool schema into context. How does it pick the right ones and learn their inputs and outputs, and when would you use MCP rather than direct API calls?

    Asked in 4 interviews, including at an enterprise AI software startup · Software Engineer · 2026

  14. We're building many agents across different business functions. Design a core agent platform that lets teams ship new agents quickly and reliably. What belongs in the shared platform, and is centralizing these capabilities even the right abstraction?

    Asked in 2 interviews, including at an e-commerce scale-up · AI Engineer · 2026

  15. Design an AI agent that colleagues message in a workplace chat tool's channels and DMs. It acts through many third-party integrations but must never surface information from channels the requesting user can't access. Walk through the agent architecture.

    Asked at an enterprise AI software startup · Software Engineer · 2026

  16. Once the agent knows which API endpoints it needs, how should it execute them: one tool per endpoint, a few generic request tools, or writing and running code in a sandbox? Which tools would you expose, and what are the trade-offs?

    Asked in 4 interviews, including at an enterprise AI software startup · Software Engineer · 2026

  17. Design an agentic platform that gathers relevant information from across the internet using many different APIs and tools, handling varied data formats. How would you implement it, and what challenges would you expect to face?

    Asked in 2 interviews, including at a startup · Founding Engineer · 2026

  18. An agent has run for an hour. Tool outputs are already trimmed, but its context window is filling up and the task isn't done. Without adding a time limit, walk through how you keep it running instead of erroring out.

    Asked in 4 interviews, including at a legaltech startup · Applied AI Engineer · 2026

  19. What's your system design philosophy for building agentic workflows from scratch in high-risk domains like finance or insurance? How do you structure guidelines and guardrails so agents don't drift in the wrong direction?

    Asked at a fintech startup · Agentic AI Engineer · 2026

LLM evaluation

9 questions
  1. Which LLMOps tools have you used to trace, evaluate and monitor LLM agents in production, and what did each one actually tell you about how your agents were behaving?

    Asked in 2 interviews, including at an e-commerce scale-up · AI Engineer · 2026

  2. What does a good evaluation setup for an LLM application look like in practice? What do you measure, and how do eval results change what you build or ship?

    Asked at a legaltech startup · Applied AI Engineer · 2026

  3. For an LLM system you shipped, how did you build the evals and the golden dataset behind them: how big, how collected and labelled (predefined metrics, human labels or both), why was it hard, and what gaps made you invest in it?

    Asked in 5 interviews, including at an e-commerce scale-up · AI Engineer · 2026

  4. Give an example where you had to evaluate a model against messy data with no real ground truth. How did you approach it, and how did you convince yourself the evaluation was meaningful?

    Asked at a market research startup · Research Engineer · 2026

  5. When an LLM-based system produces numerical figures, how do you check they are actually correct? Do you define a ground-truth set, and how would you build it?

    Asked at a defence & security startup · Senior Founding Engineer · 2026

  6. You're iterating quickly on an agentic AI product. How do you evaluate each new version, know it's actually better than the last, and catch regressions in existing behaviour before it ships?

    Asked in 2 interviews, including at a defence & security startup · Founding AI Engineer · 2026

  7. Your LLM provider deprecates the model version behind a production workflow, and the replacement behaves slightly differently. How do you monitor that the system still meets its KPIs, and how do you catch and manage regressions during the upgrade?

    Asked in 3 interviews, including at an established digital media company · AI Automation Engineer · 2026

  8. How would you judge an LLM agent's real-world performance when public benchmarks may mostly reward recall of training data? How can you tell whether a task is genuinely outside the model's training distribution?

    Asked in 2 interviews, including at a deeptech startup · ML Research Engineer · 2025

  9. How would you measure the uncertainty of a foundation model's output? Compare black-box approaches, such as perturbing the prompt and measuring output variance, with white-box approaches that use open-weight model internals, such as token probabilities or hidden states.

    Asked in 3 interviews, including at a market research startup · Research Engineer · 2026

Prompting & structured output

1 question
  1. You need to order a hotel's 40–80 photos for its gallery. Can a vision-language model actually rank that many images in one pass, or is it only reliable at labeling and captioning? How would you design it?

    Asked in 2 interviews, including at a travel tech startup · Product Engineer · 2026

LLM serving & cost

2 questions
  1. An LLM-backed processing pipeline is now in production. What would your monitoring and observability stack look like: what would you log, which health metrics would you track and alert on, and what is specific to the LLM calls?

    Asked in 3 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026

  2. Your pipeline depends on hosted vision-language model APIs. How would you handle rate limits and provider outages when 1,000 jobs arrive at once and must be processed in near real time, so queue back-pressure isn't acceptable?

    Asked in 2 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026

AI safety & security

7 questions
  1. An LLM classifier builds its user prompt as 'Classify this document:' followed by the raw document text. What security problem does this create, and how would you mitigate it?

    Asked at an AI scale-up · Applied AI Engineer · 2026

  2. A team worries that sending proprietary research data to a third-party hosted LLM could leak it. How could exposure actually happen with a hosted model versus a self-hosted one, and how would you reduce the risk?

    Asked at a biotech startup · Contract AI Engineer · 2026

  3. In AI security, where do you draw the boundary between reliability and safety issues like hallucinations and true security issues like data protection and adversarial attacks? How would you prioritise them for a healthcare deployment?

    Asked at a cybersecurity startup · Founding AI Engineer · 2026

  4. What risks do you see in letting an AI system produce financial reporting and billing, for example connected directly to an ad platform? What guardrails would you add, and how would you validate its outputs before trusting it?

    Asked at an established digital media company · AI Automation Engineer · 2026

  5. One generalist agent serves users from different teams, each entitled to different datasets. Would you split it into specialist agents? How do you enforce per-user data authorization when the agent calls tools, and where does that layer live?

    Asked in 5 interviews, including at an e-commerce scale-up · AI Engineer · 2026

  6. If you deploy agents that act inside your customers' own enterprise systems, how would you design for governance and auditability in production?

    Asked at an enterprise AI software company · Applied AI Engineer · 2026

  7. When securing LLM agents in an enterprise, where do you draw the line between what can be formally verified and what can't? How would you specify a security or context-handling property precisely enough to verify it formally?

    Asked at a cybersecurity startup · Founding AI Engineer · 2026

ML fundamentals

3 questions
  1. If every new evaluation example requires an expensive real-world experiment, how do you decide how many data points you need to judge a model's performance with statistical confidence?

    Asked at a deeptech startup · ML Research Engineer · 2025

  2. If a vision model scores hotel photos, will it generalize across very different properties, such as a small boutique hotel with few amenities versus a large resort with hundreds of photos? How would you evaluate and handle that?

    Asked at a travel tech startup · Product Engineer · 2026

  3. You propose adding a Bayesian prior to a model. How does that actually work, and why would it beat plain fine-tuning, which could learn the same structure directly in its weights?

    Asked at a startup · Research Engineer · 2026

Model training & fine-tuning

1 question
  1. Your system needs to understand images. Do you fine-tune current open-weight models to add that capability now, or wait for stronger open-weight multimodal models? How do you weigh the tradeoff and make the call?

    Asked at a market research startup · Research Engineer · 2026

Data engineering

1 question
  1. You're building an AI root-cause-analysis tool for equipment failures across large facilities. Where does the data come from (maintenance logs, sensor/IoT telemetry), how do you handle its quality problems, and where does machine learning fit?

    Asked in 5 interviews, including at a defence & security startup · Founding AI Engineer · 2026

ML system design

6 questions
  1. For an LLM product already in production, what patterns would you use to cut its latency and to build a tight user-feedback loop that actually drives model and prompt improvements?

    Asked at a healthtech startup · AI Engineer · 2026

  2. Six months after launch, your AI agent product has many customers in production and ships regular releases. Then a problem hits production. How do you work out what's going on, get things running again, and approach supporting an AI system in production?

    Asked at a defence & security startup · Founding AI Engineer · 2026

  3. Which cross-cutting concerns would you pull into shared components across agents (cost tracking, tracing and quality monitoring, chat-channel integrations, model and effort routing, workflow orchestration), and which should stay agent-specific? Why?

    Asked in 3 interviews, including at an e-commerce scale-up · AI Engineer · 2026

  4. Screen recordings and timestamped click and keystroke logs land in blob storage. Design the services that turn them into a process graph shown in a UI: the processing stages, where a VLM or LLM fits, what you'd store, and which databases.

    Asked in 2 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026

  5. How would you represent a text-heavy scientific procedure, such as a step-by-step lab protocol, so it can be embedded, compared and used efficiently by an ML model?

    Asked in 2 interviews, including at a deeptech startup · ML Research Engineer · 2025

  6. You can afford only a few dozen physical experiments. How would you design a closed loop that starts from prior results, including known failures, and uses a model to pick each next experiment to get the best outcome on that budget?

    Asked in 2 interviews, including at a deeptech startup · ML Research Engineer · 2025

Software system design

5 questions
  1. Explain the difference between client-side and server-side rendering. What are the trade-offs, and when would you choose each?

    Asked at a foundation-model scale-up · Applied AI Engineer · 2026

  2. An LLM-backed processing pipeline handles about 100 items a day. If volume grows to 1,000 a day, what exactly breaks first, which bottlenecks would you look for, and how would you adapt the design?

    Asked in 2 interviews, including at an enterprise AI software company · Applied AI Engineer · 2026

  3. When you choose tools for an internal AI automation, why should the stack your engineering team already uses affect the decision? Which dependencies would you weigh?

    Asked at an established digital media company · AI Automation Engineer · 2026

  4. An early-stage product is built on Next.js, Supabase and GitHub. As founding engineer, would you keep that stack for what comes next or change it? What would drive your decision?

    Asked at a consumer tech startup · Founding Engineer · 2026

  5. A no-code AI agent builder covers most of a workflow, but the last 20% needs custom logic, such as bespoke filtering on a third-party data API. How do you let non-technical users build that without code and without cluttering every node with options?

    Asked in 4 interviews, including at an enterprise AI software startup · Product Engineering Lead · 2023

Backend & APIs

2 questions
  1. Your product has to integrate with customers' existing systems, often decades-old on-prem software or third-party SaaS with poor APIs. How do you plug into them reliably?

    Asked at a travel tech startup · Full-Stack Engineer · 2026

    Software EngineeringIntermediatePractise related questions on Backend & APIs
  2. You need to turn a batch LLM classification script into a scalable FastAPI service. Sketch the design: the main classes and how they interact, and the overall service architecture.

    Asked in 2 interviews, including at an AI scale-up · Applied AI Engineer · 2026

    Software EngineeringIntermediatePractise related questions on Backend & APIs

Python & coding

1 question
  1. A junior colleague's quick Python script loads documents from a JSON file, classifies each with an LLM API, and writes a CSV report. Without writing code, review it: list the issues, rank them by severity, and say which you'd fix first.

    Asked in 5 interviews, including at an AI scale-up · Applied AI Engineer · 2026

    Software EngineeringIntermediatePractise related questions on Python & coding

Databases & SQL

1 question
  1. User management and CRUD data live in Postgres, while workflow graphs live in a graph database. How would you relate the two? Would you replicate or sync data between them, and how would you keep them consistent?

    Asked at an enterprise AI software company · Applied AI Engineer · 2026

Cloud & infrastructure

1 question
  1. How would you take a system with a FastAPI backend, a React client and background processing workers to production? Would it all be one service or several, and what drives that choice?

    Asked at an enterprise AI software company · Applied AI Engineer · 2026

Product sense

5 questions
  1. You join as an internal AI engineer and are asked to help the marketing team. How do you work with stakeholders to gather requirements and find the automation opportunities worth building?

    Asked at an established digital media company · AI Automation Engineer · 2026

  2. Are good AI automation targets always hard, hard-to-spot problems, or also routine work nobody should do by hand? How would you weigh a complex task done rarely against a simple task done constantly?

    Asked in 2 interviews, including at an established digital media company · AI Automation Engineer · 2026

  3. Imagine it's two years from now and we're discussing what surprised everyone in AI. What's the most counterintuitive development you expect, one that most people are missing today?

    Asked at an HR tech & recruiting startup · Product Engineer · 2026

  4. Generative AI could clearly improve customer-facing enterprise workflows like voice support. Why aren't more large enterprises adopting it? Who inside the organization is usually saying no, and why?

    Asked at an enterprise AI software startup · Product Engineering Lead · 2023

  5. Where do you think SaaS is headed? If AI-generated, dynamic front ends become the norm, what's your thesis on which parts of today's software and product stack survive?

    Asked in 2 interviews, including at a construction tech startup · Founding AI Engineer · 2026

Technical leadership

6 questions
  1. Industry research can't examine every angle the way academia can. On a project with a product deadline, how do you decide the evidence is good enough to ship an MVP rather than keep investigating?

    Asked at a market research startup · Research Engineer · 2026

  2. You've shipped ten internal AI automations for non-technical business teams. Who owns and maintains them, what go-live and support model keeps them all running, and how does aiming for team self-sufficiency change the tools you'd build with?

    Asked in 4 interviews, including at an established digital media company · AI Automation Engineer · 2026

  3. You're brought in to add an LLM capability to an existing research platform. How would you define the architecture, ship a fast first iteration in sprint one, decide what it must validate, and plan sprint two?

    Asked at a biotech startup · Contract AI Engineer · 2026

  4. Suppose you join to build ML experimentation pipelines from scratch. What infrastructure, tooling and investment would you ask for up front so you can move very quickly?

    Asked at a healthtech startup · ML Engineer · 2026

  5. With a fixed early-stage hiring budget, how many engineers would you hire for the first team, and would you include juniors or only seniors? Walk through how you'd reason about the trade-offs.

    Asked in 2 interviews, including at a defence & security startup · Founding CTO · 2026

  6. Users of an AI workflow-automation platform can request custom extensions to its third-party integrations. If it grew to millions of users, that could mean thousands of requests. How would you prioritize them, allocate integration engineers, and scale the integrations team as usage grows?

    Asked in 3 interviews, including at an enterprise AI software startup · Product Engineering Lead · 2023