
Every long document that meets a structured extraction pipeline eventually collides with a hard ceiling: a model's context window. From BERT's 512-token limit that forced EHR pipelines to invent paragraph-then-sentence-then-token fallback rules, to modern 1M+ token flagship models whose limits still fail on truly long reports, splitting is an unavoidable architectural decision. This article compares how three concrete strategies — BERT-era paragraph-aware splitting, Google's LangExtract multi-pass chunking, and Proxy-Pointer's structure-aware tree — preserve or lose document structure during extraction. It targets engineers running open-source extraction over PDFs, contracts, clinical notes, and reports that do not fit a single prompt.

BERT and its biomedical descendants such as BioBERT impose a hard input ceiling of 512 WordPiece tokens. Anything longer is silently truncated, positional embeddings stop making sense, and self-attention cost becomes prohibitive. Clinical notes rarely respect that ceiling. Discharge summaries, radiology reports, and longitudinal patient records routinely span thousands of tokens, so structured extraction over EHR text was forced to become a splitting problem long before it could become a prompting problem.
The canonical solution, documented in open EHR relation-extraction work, is a three-tier fallback rule that operates on progressively weaker structural boundaries (smitkiri/ehr-relation-extraction):
This hierarchy matters because each tier preserves a different kind of coherence. The paragraph tier keeps topical units intact, the sentence tier keeps propositional units intact, and only the token tier breaks the text mid-thought. In practice, well-formed clinical text almost never reaches the third tier, which is why the strategy earned its reputation as a reliable baseline.
The trade-off is what this rule cannot do. By design, it treats each chunk as an independent input, so any relation that crosses a paragraph boundary is lost to the encoder. Cross-paragraph discourse — for example, a medication mentioned in one section whose adverse event appears in another — disappears even when both spans are correctly tagged. BioBERT-based relation extraction in this project reached an F1 of 0.942 with gold entities but dropped to 0.86 in an end-to-end setup, a gap largely attributable to lost context between chunks (smitkiri/ehr-relation-extraction).
Conceptually, this paragraph-aware splitter is the ancestor of every modern chunking strategy. LangChain's TokenTextSplitter, sliding-window sentence chunkers with overlap buffers, and even today's multi-pass LLM chunkers descend from the same instinct: shrink the document to fit the window, and cut at the strongest available structural boundary when possible. The same three-tier instinct still shows up whenever a small encoder is reused inside an otherwise modern pipeline for cheap pre-tagging, routing, or filtering, because no larger model has yet made the problem of "where do I cut?" disappear.

A published benchmark on locally hosted entity/relationship extraction reported concrete context sizes for the three reference models in this comparison: GPT-4 at 8,192 tokens, GPT-3.5-turbo-16k at 16,385 tokens, and quantized Mistral-Instruct 7B at roughly 8,192 tokens (Kineviz). In that same study, a full prompt with few-shot examples plus the source text came in at just over 4,000 tokens — comfortably under every model's ceiling at the time, with the remaining budget reserved for the model's response.
Today's flagship numbers look very different. Gemini 2.5 Pro advertises 1,048,576 tokens, Claude 3.5 Sonnet and Claude Opus 4 reach 200,000, and GPT-4o sits at 128,000 (Machine Learning Plus). On paper, that is roughly a 100× expansion on the BERT-era 512-token ceiling and a 12× expansion on the GPT-4 8K window.
A larger budget raises the threshold for what fits in one prompt, but it does not retire chunking. Practical forces still push extraction pipelines toward split-then-extract execution:
Industry guidance echoes the same caveat: just because a window accepts a million tokens does not mean a request should fill one (Redis). The ceiling has moved up by orders of magnitude, but the architectural decision — split first, then extract — is unchanged.

lx.extract()LangExtract exposes long-document handling as three orthogonal parameters on a single call. The repository's own example, extracting characters and emotions from a full Project Gutenberg novel, sets all three at once:
result = lx.extract(
text_or_documents="https://www.gutenberg.org/files/1513/1513-0.txt",
prompt_description=prompt,
examples=examples,
model_id="gemini-2.5-flash",
extraction_passes=3, # Improves recall through multiple passes
max_workers=20, # Parallel processing for speed
max_char_buffer=1000 # Smaller contexts for better accuracy
)
Each knob addresses a different failure mode of single-prompt extraction.
max_char_buffer: chunk size and per-call accuracymax_char_buffer controls how many characters of source text are packed into each model call. The repo default of 1000 is a deliberate trade-off: smaller buffers reduce the "needle-in-a-haystack" effect, so the model is more likely to notice and faithfully quote entities that match the few-shot examples. The cost is multiplicative — halving the buffer roughly doubles the number of LLM calls, the number of prompts billed, and the amount of cross-chunk overlap you have to manage.
extraction_passes: trading latency for recallextraction_passes reruns the same prompt over each chunk N times. The first pass typically catches the prominent entities; subsequent passes surface mentions that the model skipped because they were mid-sentence, ambiguous, or sat near a chunk boundary. According to the LangExtract README, "multiple passes for higher recall" is the documented mechanism for this. The cost is linear in N: more tokens consumed, longer wall-clock per chunk, and more duplicate candidates that downstream code must deduplicate. Entities whose text genuinely straddles a chunk boundary still need an overlap buffer or a final merge pass — multi-pass alone does not glue chunks together.
max_workers: parallel calls, unchanged token billmax_workers fans chunk extraction out across a thread or process pool. With 20 workers, the same novel finishes in roughly the time of the slowest chunk rather than the sum of all chunks, so wall-clock time drops sharply. The token cost is untouched, however — every worker still bills its share of input and output tokens, and rate limits on the model API can become the new bottleneck well before CPU does.
Every LangExtract extraction carries a char_interval pointing back into the original document, which is what powers the interactive HTML visualizer and the char_interval is None filter for hallucinated outputs. That span is only trustworthy if the text passed to the model is byte-identical to the text it was indexed against. If you rewrite, reformat, or re-encode the document between passes, the offsets drift and grounding silently breaks. The safe pattern is to split once, cache the chunks, and run every pass and worker against the same cached slices.

Proxy-Pointer reframes the splitting problem. Instead of slicing a document by token count, it treats the source as a tree of self-contained semantic blocks derived from its native organization: markdown headers, contract section boundaries, and table regions. The central claim is architectural rather than empirical: an LLM is more likely to extract entities and relations correctly from one well-bounded section than from a blind hundred-page slice, because the section's context is encapsulated rather than cut. Hallucination drops, and the model no longer has to infer which clauses a stranded sentence belongs to.
Proxy-Pointer composes five engineering techniques that add no inference cost on top of ordinary chunking:
Contract > Section 7. Covenants > (c) Payment of Obligations. The LLM sees exactly where it is in the document without the chunk itself needing to contain the full parent text.Together these techniques preserve the relations that naive splitting destroys: heading-to-paragraph binding, table-to-caption binding, and the implicit scope of defined terms.
Proxy-Pointer is paired with a predictive metric called Graphability Indexing, which classifies every section of a given document family into very high, high, medium, low, and very low tiers. The rating is driven by Relational Density — the volume of actionable business edges relative to section size — not raw entity counts. That distinction prevents generic but entity-dense sections like "Notices" or "Exhibits" from being misclassified as high yield. In credit agreements, sections like "Payment of obligations" rate very high, while "Duties of Agent" or "Governing law" rate low. An important exception is preserved for ontological anchors such as "Subsidiaries," which are pinned as very high even when their edge count is small, because those edges define the corporate hierarchy that the rest of the contract's rules inherit.
The routing effect is concrete. In three large credit agreements — Emerson Electric (~228k characters), AT&T (~214k), and Texas Roadhouse (~434k) — the index let the pipeline bypass Low and Very Low sections entirely, yielding measured payload reductions of 16.10%, 33.94%, and 38% respectively, without measurable loss in extracted relations. Sections not yet in the index are mandatorily scanned, reviewed by a human, and folded back into the index, so the heatmap stabilizes after only a few documents per family. Document structure, in other words, is an accurate predictor of where the relational gold sits. (Source)

| Dimension | BERT paragraph-then-sentence | LangExtract multi-pass | Proxy-Pointer tree |
|---|---|---|---|
| Paragraph coherence | Preserved: greedy paragraph packing inside the 512-token budget, with sentence-level and finally token-level fallback (ehr-relation-extraction) | Not preserved: chunks are carved by max_char_buffer character windows, not paragraph breaks | Preserved inside each section block; paragraphs are kept intact because the skeleton tree only descends when a block overflows |
| Sentence coherence | Preserved: a complete sentence boundary is preferred over a token cut whenever one fits | Weakly preserved: smaller max_char_buffer improves per-chunk recall but can still split mid-sentence | Preserved: structure-guided chunking does not cut inside a paragraph unless it exceeds the model budget |
| Cross-section context | Destroyed: no concept of section headers, breadcrumb, or document tree — once a paragraph exits the window, the next chunk has no link to it | Destroyed by default, partially recoverable: each extraction_passes re-runs the document but cross-chunk relations only survive if a post-merge step stitches entities together (google/langextract) | Preserved: breadcrumb injection and pointer-based context encode parent section path into every chunk (Proxy-Pointer) |
| Table / figure integrity | Destroyed: tables and figures have no special handling; rows are split wherever the token counter overflows | Destroyed: character-windowed buffers slice through rows and cells | Preserved: table and figure regions are first-class nodes in the skeleton tree, kept whole unless they individually exceed the limit |
| Source-span traceability | Weak: token-level spans exist only inside each chunk; reconstructing global offsets requires bookkeeping | Strong: every extraction carries a char_interval pointing back into the source text, so ungrounded results can be filtered out (google/langextract) | Strong: each block's character span plus its breadcrumb path gives both global location and structural provenance |
max_char_buffer, extraction_passes, max_workers) trades accuracy for speed, and parallel-friendly chunking scales to novels and reports. Cost grows when you add a custom post-merge step to rebuild cross-chunk relations (google/langextract).
The right splitting strategy is rarely universal; it tracks the document's signal-to-noise profile.
Long, low-signal financial filings (≈200 pages of 10-K text). Use LangExtract with max_char_buffer=1000, extraction_passes=2–3, and enough max_workers to keep wall-clock time bounded (10–20 is a reasonable starting point for cloud endpoints). The smaller buffer forces the model to focus on roughly one table or paragraph at a time, and the multiple passes raise recall on rare entities that a single pass can miss. Merge results across passes by entity class (e.g., union on risk_factor, financial_metric) and de-conflict by char_interval so spurious extractions that cannot be grounded in the source are dropped automatically.
Contracts and credit agreements with explicit section headers. Use Proxy-Pointer's structure-aware tree plus Graphability Indexing. Build the document tree from headings, classify each section by relational density, and route only High/Medium-yield nodes (e.g., "Payment of Obligations", "Covenants", "Letters of Credit") to the LLM. Low and Very Low sections — typically "Notices", "Exhibits", "Governing Law" boilerplate — are bypassed entirely. In published runs this routing cut LLM payload by 16% on an Emerson credit agreement and up to ~38% on a Texas Roadhouse agreement, with zero section mismatches after a brief calibration pass on the first two documents.
Clinical-EHR pipelines on a small encoder. When the model is still a 512-token BERT-class encoder, the paragraph-then-sentence-then-token fallback remains the cheapest pre-tagger. Run the encoder first to mark candidate spans, then forward only spans where the encoder's confidence is below a threshold to a downstream LLM. This keeps token spend proportional to ambiguity rather than to corpus size.
The numbers you size max_char_buffer against are version-sensitive:
Splitting is a model-architecture decision, not a preprocessing afterthought. The buffer size, pass count, and section boundaries you choose should be justified by which tokens the model can see, which signals it tends to lose, and which sections of the document are worth paying for in the first place. Pick the strategy that matches the document's structure, then tune the numbers to match the model's window.