Introduction
A single model call can pull fields out of a short, tidy document. Long documents full of tables are a different job. Tables span pages, rows nest under parent rows, and every value has to pass validation and point back to where it came from. That needs structure around the model, not a bigger prompt.
LangGraph gives you that structure. It is a library for stateful, graph-based workflows: nodes are functions or model calls, edges are the transitions between them, and the graph handles retries, checkpoints and human review steps.
This post walks through the patterns I use. They come from a 12-node pipeline I built for an Australian state government client, which extracts regulatory compliance documents one table row at a time. On one named benchmark it measured 90% F1; larger document sets varied. The code below is a generic sketch, not that system's code.
Why a state machine
A document pipeline moves through defined states: parsed, extracted, validated, reviewed, saved. The transitions between them are where the interesting logic lives, so I want them written down where I can read them.
Before LangGraph I wrote this as plain Python: one orchestrator function calling the others, passing dictionaries around and handling retries by hand. Adding a step meant editing the orchestrator, and nothing recorded where a run had got to when it failed.
In LangGraph you declare the state, the nodes and the edges, and the graph runs them. Here is the skeleton:
from typing import TypedDict
from langgraph.graph import END, StateGraph
class PipelineState(TypedDict):
document_path: str
rows: list[dict] # parsed table rows, each with its page and cell position
records: list[dict] # one extraction result per row
review_queue: list[dict] # rows a person has to settle
def parse_node(state: PipelineState) -> dict:
# Deterministic: read the table structure and record where every cell sits.
return {"rows": parse_tables(state["document_path"])}
def extract_node(state: PipelineState) -> dict:
# One model call per row. A bad response costs one row, not the batch.
return {"records": [extract_row(row) for row in state["rows"]]}
def validate_node(state: PipelineState) -> dict:
ok, review = [], []
for result in state["records"]:
if result["error"]:
result = auto_correct(result) # fixable errors, each correction recorded
(review if result["error"] else ok).append(result)
return {"records": ok, "review_queue": review}
def route_after_validation(state: PipelineState) -> str:
return "review" if state["review_queue"] else "persist"
builder = StateGraph(PipelineState)
builder.add_node("parse", parse_node)
builder.add_node("extract", extract_node)
builder.add_node("validate", validate_node)
builder.add_node("review", review_node)
builder.add_node("persist", persist_node)
builder.set_entry_point("parse")
builder.add_edge("parse", "extract")
builder.add_edge("extract", "validate")
builder.add_conditional_edges(
"validate", route_after_validation, {"review": "review", "persist": "persist"}
)
builder.add_edge("review", "persist")
builder.add_edge("persist", END)
graph = builder.compile(checkpointer=checkpointer)Two things matter here. route_after_validation makes the escalation rule explicit: you can read the graph and see what happens to a row that fails. And the checkpointer saves state after every node, so a failed run restarts from the last good step instead of from the top.
One model call per row
The parse runs first and involves no model. It reads the table structure and keeps the page and position of every cell. Only then does a model see the text, one row at a time, with the parent row passed in for context.
from pydantic import BaseModel, ValidationError
class SourceCell(BaseModel):
page: int
table: int
row: int
class RowRecord(BaseModel):
group: str
item: str
value: float | None
source: SourceCell
def extract_row(row: dict) -> dict:
"""One model call for one table row."""
raw = call_model(build_row_prompt(row["cells"], parent=row.get("parent")))
try:
# The source position comes from the parser, not from the model.
record = RowRecord.model_validate({**raw, "source": row["source"]})
return {"record": record, "row": row, "error": None}
except ValidationError as err:
return {"record": None, "row": row, "error": str(err)}One call per row is slower and costs more than sending the whole document at once. What it buys is containment. A malformed response costs one row, and every row has its own record of what went in and what came out. When a reviewer asks why a value is wrong, there is one prompt and one response to look at.
In this sketch the source position is attached in code rather than asked of the model, so the model can't invent where a value came from.
Validate, correct, then ask a person
Every row is checked against the target schema before anything is written. Pydantic does the checking: model_validate() raises a ValidationError that names the exact fields that failed, which is far easier to act on than an import error three steps later.
Some failures can be fixed without a person, such as a value in the wrong format. The correction step fixes those and records each change, so the fix is visible later. Rows that still fail go to a review queue. On the review screen a person opens the source page with the cell highlighted, settles the row and approves the records before they are exported.
What I tried first
My first prototype sent the whole document to a model and asked for JSON in one pass. Long documents overflowed the context window and the output was silently truncated. The model invented values where a table was ambiguous or split across pages. And there was no intermediate state, so I couldn't tell where an error came from.
Batching 10 to 20 rows per call came next. It was cheaper, but one malformed response failed the whole batch.
An earlier version of this post taught a consensus pattern: run several model providers in parallel and merge their answers by weighted vote. That is not how the pipeline works. Its own earlier design ran two PDF parsers side by side, Docling and MinerU, and merged their output, marking values where both agreed. Both are superseded by the deterministic parse and the per-row call above.
Streaming progress to the review screen
Document pipelines run for minutes. One 190-page production document took about 31 minutes. If there is a web screen on top, it needs progress as it happens, or people watch a spinner and assume something broke.
LangGraph can stream graph events with astream_events, and FastAPI can pass them to the browser as Server-Sent Events:
import json
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
app = FastAPI()
async def stream_pipeline(document_path: str):
"""Stream pipeline events as SSE."""
initial_state = {
"document_path": document_path,
"rows": [],
"records": [],
"review_queue": [],
}
async for event in graph.astream_events(initial_state, version="v2"):
name = event.get("name", "unknown")
if event["event"] == "on_chain_start":
yield f"data: {json.dumps({'status': 'node_start', 'node': name})}\n\n"
elif event["event"] == "on_chain_end":
yield f"data: {json.dumps({'status': 'node_complete', 'node': name})}\n\n"
yield f"data: {json.dumps({'status': 'complete'})}\n\n"
@app.post("/process-document")
async def process_document(document_path: str):
return StreamingResponse(
stream_pipeline(document_path),
media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)On the front end, EventSource (or fetch with a ReadableStream) consumes the stream. In the pipeline the case study describes, rows appear in the review grid as they are extracted, so a reviewer can start before the run finishes.
Lessons
Carry errors in the state. Every node that calls a model or reads a file can fail. Put the error on the row it belongs to, as the sketch does, and route on it. A run that keeps going and parks the failures is worth more than one that stops on the first bad row.
Trace every call. I use Langfuse for pipeline traces and Logfire for inspecting Pydantic validation. When a value comes out wrong, I want the exact prompt, the exact response and the validation result for that row, not a guess.
Validate at every boundary. Check model output with Pydantic before it touches anything downstream, and check it again against the target system's rules before export.
Keep the source with the value. For every extracted value, store the page and cell it came from and the call that produced it. When a reviewer or an auditor asks why a value is there, "the model said so" is not an answer.
Respect rate limits. One call per row adds up fast across a batch of documents. Put a semaphore or token bucket in front of the model client, and watch queue depth and throughput. A pipeline that hits a rate limit and silently drops rows is worse than no pipeline.
Conclusion
The state machine makes the pipeline readable, the checkpoints make failures cheap, and validation before every write makes the output something a person can check. The model does one small job per row, inside a structure that decides what happens when it gets that job wrong.
The Document pipeline case study shows these patterns in the system I built. For a different use of the same ideas, code review with a person approving every fix, see Sentinel: code review with approval gates.
