Data Readiness Checker for GenAI

A free 30-question assessment of whether your data platform is ready to support enterprise GenAI at production quality — across catalog & lineage, data quality, access & permissions, retrieval & vector, and compliance & governance. About 10 minutes. Nothing submitted.

30 questions · max score 90 · operational counterpart to the Data & Infrastructure framework dimension

Progress
0 / 30

Catalog & Lineage

Is your data findable and traceable — with clear ownership and downstream lineage?

  • 1. How does someone in your org find a dataset relevant to a new use case?
  • 2. What percentage of production datasets have a documented owner?
  • 3. Can you trace a downstream metric back to its source tables and transformations?
  • 4. How is training-corpus provenance tracked for GenAI use cases?
  • 5. How stable are dataset schemas?
  • 6. How is data-source deprecation handled?

Data Quality

Do you know when data breaks, and can you trust what you retrieve?

  • 1. How do you detect data-quality issues in critical pipelines?
  • 2. How are freshness SLAs defined and monitored?
  • 3. How do you handle duplicate / conflicting records across sources?
  • 4. How is metric definition drift prevented?
  • 5. For RAG systems, how do you measure retrieval quality?
  • 6. How is training / eval data curated for quality?

Access & Permissions

Can the right people (and agents) reach the right data without a ticket?

  • 1. How does a new team get access to relevant datasets?
  • 2. How is access to sensitive data (PII, financials, HR) controlled?
  • 3. How do LLM agents authenticate when reading data on a user's behalf?
  • 4. How are access reviews conducted?
  • 5. How is production data used for testing / development?
  • 6. How is third-party (vendor, partner) data access governed?

Retrieval & Vector

Is retrieval infrastructure (vector store, hybrid search) in a shape that supports production RAG?

  • 1. Where do your RAG systems store embeddings?
  • 2. How are chunks and embeddings kept fresh as source documents change?
  • 3. What retrieval strategy do your RAG systems use?
  • 4. How do access controls carry from source data into the vector store?
  • 5. How is prompt caching used in retrieval-heavy workloads?
  • 6. How do you handle document versioning in retrieval?

Compliance & Governance

Do PII redaction, retention, cross-border, and DSR flows meet regulatory expectations?

  • 1. Where does PII redaction happen in your GenAI pipelines?
  • 2. How are data-subject deletion requests handled for GenAI-adjacent data (prompts, completions, embeddings)?
  • 3. How is cross-border data flow controlled for GenAI systems?
  • 4. How is retention defined for prompt / completion logs?
  • 5. For any high-risk AI system, do you have training / eval data documentation ready for regulatory audit?
  • 6. How is copyright / licensing status verified for training / retrieval corpora?

FAQs

  • How long does the assessment take?

    About 10 minutes. Thirty multiple-choice questions across five sub-dimensions of data readiness for GenAI.

  • What are the five sub-dimensions?

    Catalog & Lineage, Data Quality, Access & Permissions, Retrieval & Vector, Compliance & Governance. Each gets 6 questions and a per-sub-dimension score.

  • How does this relate to the GenAI Maturity Framework?

    This is the operational counterpart to the Data & Infrastructure dimension of the framework. Data readiness is one of the most common blockers to scaling GenAI beyond pilots.

  • What if we do not use RAG?

    The catalog / quality / access / compliance dimensions still apply. If a specific retrieval-focused question doesn't apply, choose the lowest score to represent the absence — the tool interprets that correctly.

  • Are my answers stored anywhere?

    No. Nothing is submitted or logged. The quiz runs entirely in your browser; answers persist to localStorage on your device only.

Framework dimensions

Next steps