FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist · 4 minute read

AI Data Readiness Checklist

Data is ready for an AI use case when every source it needs is programmatically accessible with change capture, parsed with structure preserved, enriched with metadata, mapped to enforceable permissions, of known quality and freshness with authoritative sources designated, governed for sensitivity and retention, and complemented by written rules for the tacit knowledge the system must apply.

By FISTA Solutions· AI-Native Engineering Team·
AI Data Readiness Checklist article cover

Data readiness is the best single predictor of whether an AI project reaches production, and the most common line missing from its business case. This checklist scores the sources a specific use case needs across seven dimensions, so gaps are found before they become production failures and remediation is priced honestly. It is the operational form of the data readiness for generative AI whitepaper and complements ai data readiness and the broader ai readiness assessment.

Who should use this checklist?

Engineering teams starting an AI use case, source owners whose content or data will be used, security teams responsible for access control, and business owners who need an honest estimate of the work.

Have you scoped the sources to the use case?

  1. The use case is defined and its required sources listed.
  2. Each source has a named owner.
  3. Sources not needed for the first increment are explicitly deferred.
  4. Historical data for evaluation is identified.

Is each source accessible?

  1. Content or data can be retrieved programmatically through a supported connector or API.
  2. Incremental updates and deletions are detectable.
  3. Authentication and authorization for the connector are arranged.
  4. Rate limits and export constraints are known.
  5. Sources that only support manual export are flagged as gaps.

Does parsing preserve structure?

  1. Parsers per format (documents, slides, PDFs, scans, pages, spreadsheets) are identified.
  2. Headings, tables, lists, and relationships survive parsing, verified on samples.
  3. Scanned or image content has an OCR path.
  4. Chunking strategy is structure-aware and testable.
  5. Structured sources have documented schemas.

Reference: how to build an ocr pipeline with llms and what is chunking in rag.

Is metadata sufficient?

MetadataPresent?Why it matters
Source and locationCitation and audit
Author and ownerAuthority and maintenance
Created, modified, effective datesFreshness and versioning
Document or record typeFiltering and routing
Business unit, region, audienceScoping and permissions
Sensitivity labelHandling rules
VersionSuperseded content detection
Timestamps on structured recordsPoint-in-time correctness

Are permissions known and enforceable?

  1. Each source exposes access-control information per item.
  2. The permission model is consistent with the identity provider's groups and roles.
  3. Inherited and broken inheritance, sharing links, and broad-access folders are identified.
  4. Entitlements can be resolved at query time for a user.
  5. Permission cleanup needed in the source is scoped and owned.
  6. Compliance tests with restricted items and users are planned.

Reference: ai access control.

Are quality and freshness known?

  1. Content has been sampled for accuracy, currency, duplication, and contradiction.
  2. Authoritative sources per topic are designated.
  3. Superseded and stale content is identified for archiving or exclusion.
  4. Duplicates are identified for deduplication.
  5. A refresh strategy per source matches its change rate.
  6. Structured data has quality checks for nulls, ranges, and consistency.

Reference: ai training data checklist.

Is governance settled?

  1. Each source is classified by sensitivity with handling rules.
  2. Retention rules for source content, prompts, outputs, and traces are defined.
  3. Residency requirements are known and satisfiable.
  4. Provider terms are consistent with the data's regulation; redaction needs are identified.
  5. Usage policy confirms the content may be used for this AI purpose.
  6. Lineage from source to output can be maintained.

Reference: ai data governance and ai data privacy compliance.

Is tacit knowledge captured?

  1. The rules and exceptions experts apply are identified as written or unwritten.
  2. Discovery (shadowing, case walkthroughs, interviews) is planned to write them down.
  3. Written rules become content for retrieval and cases for evaluation.
  4. Experts are allocated to label evaluation data.

Reference: how to run ai discovery and the spec-driven development for AI whitepaper.

Have you scored and sequenced remediation?

  1. Each source has a score per dimension with evidence.
  2. Remediation per gap is classified: engineering, data work, organizational, governance.
  3. Effort and owners are estimated per remediation.
  4. Remediation is sequenced so the first increment can proceed.
  5. The business case reflects the remediation cost.

Reference: the AI total cost of ownership whitepaper.

How should the results be used?

A readiness report with scores, gaps, remediation plan, and cost is the deliverable. Use cases whose sources score poorly on access or permissions should start with remediation rather than model work. Use cases with strong scores can proceed to specification and evaluation dataset construction immediately.

How FISTA Solutions uses this checklist

FISTA Solutions runs this assessment at the start of every generative AI and predictive engagement, scoped to the first workflow, with source owners, security, and domain experts. Forward deployed engineers run discovery to capture tacit knowledge, the AI enablement practice builds the connectors, parsing, metadata, and permission-aware retrieval that close the gaps, and AI agents are deployed only on sources that meet the bar. The record behind the approach is 150+ projects with 99.9% uptime.

To run a data readiness assessment, message FISTA on WhatsApp, or read how to prepare data for ai for the remediation practices.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How do you assess data readiness for AI?

Identify the sources a specific use case needs, score each on access, structure, metadata, permissions, quality and freshness, and governance, identify tacit knowledge that must be written down, estimate remediation per gap, and sequence it so the first increment can proceed.

02What is the most common data readiness gap?

Permissions: sources with inconsistent, inherited, or broken access control that cannot be enforced at retrieval without cleanup. Close behind are missing metadata such as effective dates and ownership, and contradictory or stale content.

03Does data readiness apply to structured data too?

Yes. Structured sources need schema documentation, quality checks, timestamps for point-in-time correctness, lineage, and access control, alongside the unstructured content checks. Predictive use cases lean on the structured items; generative use cases on both.

04How long does remediation take?

It depends on source count, parsing difficulty, metadata gaps, and permission consistency. Scoping to the first use case keeps it tractable, and the assessment produces the estimate. Skipping the assessment does not avoid the work; it delays it to production.

05Who should complete this checklist?

The engineering team with the source owners, the security team for permissions, and the domain experts for quality, authority, and tacit knowledge. Readiness is a joint finding, not an engineering opinion.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project