Skip to content

How iCEV Automated Hundreds of Hours of Data Extraction with AI

  • by

Overview

iCEV provides educational content and resources that help learners prepare for academic and career success. An important part of that work is mapping individual lessons to applicable academic standards.

These lesson-to-standard correlations support the creation of alignment documents for educators and institutions. They also provide historical data that iCEV can use to improve its AI-assisted alignment processes.

Over time, iCEV accumulated hundreds of Excel and PDF files containing this information. Extracting and organizing the data manually would have required hundreds of hours, so iCEV partnered with IntelliTect to develop a faster, scalable approach.

Challenge

The alignment data was distributed across files with widely different structures and formats. Some contained hundreds of rows, while others stored important details in freeform text that combined lesson names with references such as presentation slide numbers.

A traditional extraction script could not reliably account for all these variations. Manually reviewing every file would have been slow, costly, and difficult to repeat as iCEV’s content library continued to grow.

Any automated solution also needed to meet a high standard for reliability. The extracted data would eventually support downstream AI processes, so plausible-looking but incorrect results could not simply be accepted without verification.

Large files introduced another complication. Some contained more information than an AI model could process in a single request, requiring a way to divide the content without losing the hierarchy and context needed to interpret it correctly.

Solution

IntelliTect developed a C# and .NET application that uses Microsoft Foundry and OpenAI models to extract lesson names and standards correlations from Excel and PDF files.
The solution combined AI interpretation with conventional software validation to address the limitations of either approach on its own.

Adapting to inconsistent files

Rather than relying on a fixed template, the application used AI models to interpret unstructured and semi-structured content. This allowed it to identify lessons, standards, and relationships across files with different layouts and naming conventions.

Preserving context in large files

For files that exceeded model context limits, IntelliTect implemented an intelligent chunking process that retained the relevant hierarchy as the content was divided.
The application could then process multiple sections in parallel, improving performance without separating individual records from the context needed to understand them.

Verifying AI-generated results

Because AI output can contain inaccuracies, IntelliTect built verification directly into the extraction workflow.

The application asked the model to return both the extracted lesson and its location in the source file. Deterministic code then checked whether the lesson appeared at the cited location. When a result could not be verified, it was automatically returned to the model for correction.

An additional model-based review checked the extracted data for consistency. Together, these safeguards reduced the risk of incorrect information reaching iCEV’s database.

Preparing the data for downstream use

The application transformed the extracted information into a consistent, structured format designed for database ingestion. This gave iCEV a practical way to make its historical alignment data available to the systems that support alignment analysis and document creation.

IntelliTect completed the application in 56 hours over two weeks. The engagement drew on IntelliTect’s Microsoft specializations in AI Platform and AI Apps, as well as the expertise of AI Engineer Associate-certified software architect Clayton Gravatt.

Outcome

The completed application processed hundreds of files that would otherwise have required hundreds of hours of manual extraction.

Instead of asking employees to review files row by row, iCEV received structured output ready for ingestion into its database. The combination of AI extraction, source citations, and deterministic validation also gave the company greater confidence in the information being prepared for its downstream AI systems.

The application provides several lasting benefits:

Less manual work: Hundreds of files can be processed without requiring employees to extract each correlation individually.

More trustworthy results: Model-based review and source-level validation help identify and correct questionable output.

A scalable process: iCEV can use the application to process additional files as its content library and standards coverage grow.

More time for educational work: Team members can focus on content development and other higher-value initiatives instead of repetitive data entry.

By pairing the flexibility of AI with the reliability of deterministic software, IntelliTect helped iCEV turn a large collection of inconsistent source files into structured data that can support the next stage of its educational technology platform.