# Claude API OCR on Financial-Aid Letters: Check the Math, Not the Reading

- Author: Abdullah Chaudary
- Category: Self-Hosted AI
- Published: 2026-10-09T21:37:01.654Z
- Updated: 2026-10-09T21:37:01.839Z
- Reading time: 8 min
- Tags: Claude API, OCR, Document AI, LLM Evaluation, FastAPI

> LLM OCR on financial-aid letters: a strict schema, deterministic normalising, sums that must reconcile, and flags for people. Reading the page is only the first check.

**TL;DR:** LLM OCR document extraction is only trustworthy when you evaluate the extraction, not the reading. On a platform that compares college financial-aid letters, Claude reads the page, but the answer comes from what happens next: a strict schema of the fields that matter, rules that normalise each school's terms and periods, sums that must reconcile to the net price, a report that shows its math, and flags that send anything unusual to a person.

## Why Is Reading the Page Not Enough for Financial-Aid Letters?

Reading the page is not enough because a financial-aid letter can be read perfectly and still produce the wrong numbers. Every school uses its own layout, its own names for the same aid, and its own periods. A model that transcribes "Federal Direct Unsubsidized" correctly has still failed if that line ends up counted as free money.

This is not a rare edge case. In December 2022 the U.S. Government Accountability Office reviewed more than 500 offers from a nationally representative sample of 176 colleges and found that "about 91% of colleges understate or don't include the net price in their offers," as [CAPPS' reprint of GAO's highlights](https://cappsonline.org/financial-aid-offers-action-needed-to-improve-information-on-college-costs-and-student-aid/) puts it. No college in the sample followed all ten of GAO's best practices.

The voluntary [College Cost Transparency standard](https://www.aplu.org/wp-content/uploads/CCTI-Principles-and-Standards-Final.pdf), which [more than 720 colleges have joined](https://www.reed.edu/newsroom/press-releases/2026/reed-college-joins-national-college-cost-transparency-initiative.html), defines the number families actually need: "An estimated net price for the student, derived by subtracting grants and scholarships from the total Cost of Attendance." Loans and work-study are not subtracted. A comparison tool lives or dies on that one rule.

## What Does the Pipeline Look Like From Upload to Report?

The pipeline runs upload, OCR, extraction, comparison and report, with a check at every boundary. I built it on [an AI document-evaluation platform for an EdTech and finance client](/about), where I owned the proposal and client onboarding through to shipping the product. The main app is Vue.js and FastAPI, with the Claude API doing the OCR and evaluation.

1. **Upload.** Families upload their letters in whatever format they have: PDFs, scans, and sometimes a phone screenshot. The Vue front end compresses images in the browser before upload.
2. **Redaction.** Names and sensitive details arrive redacted, and Claude only processes the fields the schema asks for.
3. **OCR and extraction.** FastAPI sends the page to Claude, which reads it and fills the schema.
4. **Comparison.** Code, not the model, normalises every school's figures onto the same terms and computes the comparison.
5. **Report.** The app compiles a report that shows the calculation behind every number.

The first implementation of the platform ran on AWS, with models served through Amazon Bedrock and a .NET backend. The bills were high, and .NET has thin library support for this kind of document and LLM work. Moving the pipeline to Python and FastAPI put it next to the best-supported tooling, and moving off AWS is a decision I have [written up separately](/blog/aws-self-hosted-migration-cost).

## Which Fields Should an LLM Extract From an Aid Letter?

Extract every line that changes what a family pays: tuition and fees, housing and meals, books and other direct and indirect costs, every grant and scholarship, every loan by type, work-study, and any conditions attached to renewal. Extract each one as printed, with its label, amount, period and page, and let code decide what it means.

The schema is the contract. Ours is a clear set of fields with rules for each, and nothing outside it gets processed. A simplified sketch of the pattern, not the client's code:

```python
from enum import Enum
from pydantic import BaseModel

class Period(str, Enum):
    year = "year"
    semester = "semester"
    quarter = "quarter"
    unknown = "unknown"

class Kind(str, Enum):
    cost = "cost"
    grant = "grant"      # grants and scholarships: the only aid that lowers net price
    loan = "loan"
    work = "work"        # work-study
    unknown = "unknown"

class LineItem(BaseModel):
    label_as_printed: str    # "Presidential Award", "Federal Direct Unsubsidized Loan"
    amount_as_printed: str   # "$4,250.00", kept as text so code parses it
    period: Period
    kind: Kind
    page: int

class AidLetter(BaseModel):
    academic_year: str
    items: list[LineItem]
    net_price_as_printed: str | None   # only if the letter prints one
```

Keeping the amount as printed matters. The model's job is to find and transcribe the figure; turning "$4,250.00 per semester" into an annual number is arithmetic, and arithmetic belongs in code.

## How Do You Get Structured JSON From the Claude API?

In 2026 the Claude API enforces the schema for you. [Structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) are generally available: the request carries the JSON schema in `output_config.format`, and the Python SDK builds it from a Pydantic model and returns the parsed result. The older `output_format` parameter and its beta header are deprecated.

```python
import anthropic

client = anthropic.Anthropic()

result = client.messages.parse(
    model=MODEL,               # pinned model id; we run Claude Sonnet 4.6
    max_tokens=4096,
    messages=[{
        "role": "user",
        "content": [
            {"type": "document", "source": {
                "type": "base64", "media_type": "application/pdf", "data": pdf_b64}},
            {"type": "text", "text": EXTRACTION_RULES},
        ],
    }],
    output_format=AidLetter,
)
letter = result.parsed_output
```

[PDF support](https://platform.claude.com/docs/en/build-with-claude/pdf-support) sends each page as extracted text plus a page image, so the model sees the table layout as well as the words. Screenshots go in as image blocks.

Three limits of the schema guarantee are worth designing around. Numeric constraints such as `minimum` and `maximum` are not enforced in the schema, so range checks live in your validation. A refusal or a response cut off at `max_tokens` can break schema compliance. And structured outputs cannot be combined with citations, which return a 400 error together, so the `page` field in the schema is how each value points back to its source.

We run Claude Sonnet 4.6 because it costs less than the larger models and its accuracy on these letters is very good.

## How Do You Normalise Amounts and Terms Across Schools?

Normalise with deterministic rules, never with the model's judgment. Every school names and splits aid differently, so a rules layer maps printed labels onto one vocabulary, converts every period to an academic year, and parses amounts as exact decimals. Anything the rules cannot place is marked unknown and flagged instead of guessed.

```python
from decimal import Decimal
from pydantic import BaseModel, model_validator

PERIODS_PER_YEAR = {"year": 1, "semester": 2, "quarter": 3}
LOAN_WORDS = ("loan", "plus", "subsidized", "unsubsidized")

def kind_of(item: LineItem) -> str:
    label = item.label_as_printed.lower()
    if any(word in label for word in LOAN_WORDS):
        return "loan"            # a loan is a loan, whatever the letter calls it
    return item.kind.value.lower()

def annual(item: LineItem) -> Decimal:
    amount = Decimal(item.amount_as_printed.replace("$", "").replace(",", ""))
    return amount * PERIODS_PER_YEAR[item.period.value]   # "unknown" raises: flag it

class Comparison(BaseModel):
    cost_of_attendance: Decimal
    grants: Decimal
    loans: Decimal
    work: Decimal
    net_price: Decimal

    @model_validator(mode="after")
    def net_price_counts_grants_only(self):
        if self.net_price != self.cost_of_attendance - self.grants:
            raise ValueError("net price must be cost of attendance minus grants and scholarships")
        return self
```

`Decimal` is there for a reason. Python's [decimal documentation](https://docs.python.org/3/library/decimal.html) shows that `0.1 + 0.1 + 0.1 - 0.3` is exactly zero in decimal and not in binary floating point, and money comparisons should never depend on rounding noise. Pydantic's [model validators](https://docs.pydantic.dev/latest/concepts/validators/) run after the whole model is built, which is where cross-field rules like the net-price check belong.

The loan rule is the one that protects families. The College Cost Transparency standard says "All loans should be unambiguously labeled as such," using the word "loan," because offers that blend loans into aid make a school look cheaper than it is. The code overrides the model whenever a label says loan.

Hard bounds catch transcription errors cheaply. For the 2026-27 award year, the [Federal Pell Grant maximum is $7,395](https://fsapartners.ed.gov/knowledge-center/library/dear-colleague-letters/2026-01-30/2026-27-federal-pell-grant-maximum-and-minimum-award-amounts) and the minimum is $740, so an annual Pell figure outside that range is an extraction error, not a generous school.

## How Do You Check That the Extraction Is Right?

Check it in layers: hand-verified letters while tuning, exact-match scoring on the fields that matter, sums that must reconcile, and a report that shows every calculation. No single layer is enough. Together they turn "the model read it" into "the numbers add up and a person can see why."

**Hand-checked letters first.** While we tuned the prompts, schema and rules, we checked the extracted fields against the letters by hand. Anthropic's [guide to building evals](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests) describes the same discipline at scale: exact-match evals "measure whether the model's output matches a predefined correct answer," and many automated cases beat a few graded by eye. Score each field separately, because a letter that is mostly right can still be wrong on the one line that decides the comparison.

**Then the math.** Every comparison is reconciled: cost of attendance minus grants and scholarships must equal the net price, the line items must add up to the totals the letter prints, and every figure must sit inside its known bounds. When a letter prints its own net price and it disagrees with ours, that difference is a finding.

**Then the report shows its work.** After the tuning phase we added "the math behind it" to the report: the in-depth calculation for every school, from each printed line to the final net price. A family, or a reviewer, can trace any number back to the line and page it came from.

## Where Does Human Review Still Belong?

Human review belongs wherever the input or the pipeline is unusual. The app flags uploads that need a person instead of trusting them: phone screenshots, documents that arrive with names or sensitive details still visible, items the rules could not classify, checks that fail, and any error anywhere in the processing pipeline.

That matches the vendors' own advice. Anthropic's [vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision) warns that Claude "might hallucinate or make mistakes when interpreting low-quality, rotated, or very small images under 200 pixels," and recommends careful review for high-stakes use. Over-compressing a screenshot in the browser can push it into exactly that territory, so compression has a floor.

The Claude API does not return a per-field confidence score, so confidence has to come from the checks. A field that fails a bound, a sum that does not reconcile or a label the rules do not know is the confidence signal, and each one is a reason to show the document to a person.

## What Would I Change in 2026?

Two things: the normalising rules and the model. Formats and terms differ from school to school, and as the number of API users grows, hand-maintained rules stop scaling. We have decided to move the rules onto a shared, standard schema, either an open-source one or a full implementation of an existing standard, so a new school is configuration rather than code.

The second is moving from the Claude API to self-hosted models. The evaluation harness is what makes that move safe: with hand-checked letters, field-level scores and reconciliation checks in place, swapping the model becomes a measurement, not a leap of faith. It is the same principle I follow for [retrieval in a self-hosted RAG stack](/blog/qdrant-bge-m3-hybrid-search-rag): measure each stage on its own.

## Limitations

- The code above is a simplified sketch of the pattern, written for this article; the client's codebase stays under NDA. The snippets have not been run as shown.
- Structured outputs guarantee the shape of the JSON, not the truth of the values. Every number still needs the checks above.
- Bounds such as the Pell maximum change every award year and must be updated with the year on the letter.

## References

- [Structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs), Claude API documentation
- [PDF support](https://platform.claude.com/docs/en/build-with-claude/pdf-support), Claude API documentation
- [Vision](https://platform.claude.com/docs/en/build-with-claude/vision), Claude API documentation
- [Define success criteria and build evaluations](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests), Claude API documentation
- [Financial Aid Offers: Action Needed to Improve Information on College Costs and Student Aid](https://cappsonline.org/financial-aid-offers-action-needed-to-improve-information-on-college-costs-and-student-aid/), GAO-23-104458 highlights, December 5, 2022, reprinted by CAPPS
- [College Cost Transparency Initiative: Principles and Standards](https://www.aplu.org/wp-content/uploads/CCTI-Principles-and-Standards-Final.pdf)
- [2026-27 Federal Pell Grant maximum and minimum award amounts](https://fsapartners.ed.gov/knowledge-center/library/dear-colleague-letters/2026-01-30/2026-27-federal-pell-grant-maximum-and-minimum-award-amounts), Federal Student Aid
- [Validators](https://docs.pydantic.dev/latest/concepts/validators/), Pydantic documentation
- [decimal: Decimal fixed-point and floating-point arithmetic](https://docs.python.org/3/library/decimal.html), Python documentation

Have you built extraction on documents where every issuer uses its own format? I would like to hear how you check the numbers, not just the reading.
