A practical guide to AI due diligence and GenAI diligence — how PE, corp dev, and M&A buyers test model quality, data rights, product risk, unit economics, and moats before banking an artificial intelligence thesis.
Many targets now claim “AI-powered” growth, margin, or defensibility. AI due diligence is the work that decides whether those claims survive contact with model cards, data contracts, eval suites, inference bills, and customer liability. It is not the same as technology due diligence (stack and scalability), product due diligence (roadmap and fit), cybersecurity diligence (security posture), or SaaS metrics diligence (ARR/NRR). AI diligence underwrites the intelligence layer: what is proprietary, what is rented, and what can break the model at scale.
| Workstream | Primary question | Typical output |
|---|---|---|
| Technology DD | Is the stack scalable, maintainable, and secure enough? | Architecture map, debt, team capacity |
| AI / ML DD | Is the intelligence real, owned, measured, and economic? | Model/data inventory, evals, cost curves, IP risk |
| Product DD | Does the product solve a durable job for buyers? | Roadmap fit, differentiation, adoption |
| Cyber / privacy DD | Can attackers or regulators stop the business? | Threat surface, controls, privacy program |
| SaaS / commercial DD | Are growth and retention real? | ARR quality, churn, win/loss |
Catalog every model in production and near-production: proprietary trained models, fine-tunes, embeddings, classical ML, third-party foundation APIs, and open-source weights. For each, document task, owner, latency SLA, fallback path, versioning, and whether it sits on the critical path of revenue or safety. Separate “AI in the deck” from “AI in the product path.” Map orchestration (agents, tools, RAG pipelines) so buyers see the full system, not a single model name.
Trace training, fine-tuning, evaluation, and retrieval corpora: customer data, partner feeds, licensed datasets, scraped web, synthetic data, and employee-created labels. Confirm contractual rights to train, improve, and commercialize; deletion and opt-out obligations; cross-border transfer limits; and whether licenses (including open-source dataset terms) conflict with the product. Assess data quality, labeling process, drift monitoring, and PII handling — and connect findings to data privacy diligence and IP diligence.
Demand reproducible evals: offline benchmarks, golden sets, human preference or expert review, online A/B or shadow tests, and task-specific metrics (not only generic leaderboard scores). For GenAI, test hallucination rate on customer-critical tasks, refusal behavior, jailbreak/prompt-injection resistance, and regression gates on model upgrades. Require model cards or equivalent documentation, known failure modes, and a change-management process when the provider ships a new base model.
Map where AI outputs reach customers, regulators, or automated actions without human review. Inventory safety incidents, customer complaints, content moderation, logging of prompts/outputs, and data leakage paths (training on tenant data, shared context windows, plugin tools). Align product claims with reality: copilots vs autonomous agents, accuracy warranties, and insurance implications (see insurance diligence). Red-team or at least sample adversarial prompts on high-risk flows.
Build cost-to-serve for inference, embedding, storage, GPU/CPU, evaluation, human-in-the-loop review, and vendor markups. Stress token or request growth vs pricing power. Test concentration risk on a single model provider, region, or GPU supplier and the switching plan (second provider, open weights, distillation). Separate gross margin impact of AI from marketing claims. For platforms selling AI features, verify packaging, usage caps, and whether AI is margin-accretive or a growth giveaway.
Decide whether defensibility is proprietary data, distribution, workflow lock-in, model performance, brand trust, or none of the above. Map key ML/product talent, documentation quality, and bus factor. Review patents, trade secrets, open-source outbound obligations, and ownership of fine-tunes and outputs. Confirm AI governance: policies, risk tiers, board/IC reporting, regulatory horizon (sector AI rules, consumer protection, sector-specific model risk). Tie people risk to people diligence and compliance to regulatory diligence.
DI20-WELCOME) — useful for triage, not a full model audit, red team, or security assessment.
| Stage | AI focus | Buyer action |
|---|---|---|
| Pre-LOI / IOI | Thesis materiality, public product claims, hiring/IP signals | Price only defensible AI value; flag data/vendor risk |
| LOI / exclusivity | Model inventory, data contracts, rough cost-to-serve | Data request list; access to ML leads; eval samples |
| Confirmatory DD | Evals, rights, liability, switching plan, talent | Red/amber/green by system; model cases; kill criteria |
| SPA / financing | IP/data reps, AI warranties, escrow of critical assets | Align definitions; financing model matches diligence |
| Close / Day-1 | Access continuity, key person, vendor accounts | No silent model swaps; logging and rollback live |
| Signal | Severity | Why it matters |
|---|---|---|
| “AI-powered” with no model inventory or evals | Deal-Killer | Thesis un-underwritable; often marketing only |
| Thin wrapper on one foundation API, no switching plan | Deal-Killer | Margin and product hostage to vendor pricing/policy |
| Training on customer data without clear rights | Deal-Killer | Contract and regulatory blow-up risk |
| Inference cost curve kills unit economics at scale | High | Growth destroys cash and gross margin |
| No regression tests when base models change | High | Silent quality cliff after provider updates |
| Key ML talent is a single person with tribal knowledge | High | Execution and maintenance single point of failure |
| Material copyright / training-data litigation exposure | High | IP and brand risk not priced |
| Autonomous actions without human review on high-risk flows | Watch | Liability and brand asymmetric downside |
| Approach | Typical cost | Timeline | Best use |
|---|---|---|---|
| Full AI / ML + safety deep dive | $40K–$250K+ | 3–10 weeks | AI-core thesis, exclusivity, IC-grade risk |
| Focused model + data rights review | $25K–$90K | 2–5 weeks | Clear product AI, limited GenAI surface |
| Public first-pass risk pack | $49 | Minutes to hours | Triage before LOI / shortlist |
Before LOI, buyers use structured public research to pressure-test whether an AI story is even plausible: product claims vs demos, hiring and research signals, patents and open-source footprint, pricing and packaging, customer logos, and competitive density of similar wrappers. After LOI, the same hypotheses drive the data-room request list — model inventory, data contracts, eval reports, cost telemetry, talent map — so advisors do not spend weeks on marketing slides. The pack is screening research, not a substitute for model audits, red teams, or legal IP opinions.
⇧ Already delivered: Tesla (TSLA) · Alphabet (GOOGL) · Palantir (PLTR) — real orders, real SEC data, every claim source-cited.
Get a structured first-pass diligence pack on your target — useful input for AI / GenAI hypotheses, not a full model audit or security red team.
Order report $39.20 → Free brief Sample PDF