Research

Which Pages of a 200-Page Pension File Actually Matter?

KleioBase EditorialSeptember 9, 202619 min read
Share

The file arrived. You had been waiting weeks for it, you paid for it, and it is 187 pages of handwriting. You opened it, read eleven pages, found a surgeon describing your great-great-grandfather's rheumatism in 1884 in unpleasant detail, and closed it again. That was in March.

This is the most common unfinished job in American genealogy and almost nobody writes about it. There is a great deal of published advice on how to order a Civil War pension file, and a great deal on what such files contain in general. There is essentially nothing on what to do when one is sitting on your desk and you have a Saturday.

So here is the useful part, which is that two decisions have already been made for you, both by the federal government, and neither has ever been written up for people who actually need them.

The short answer. Pension files are filed in a mandated order, so you can navigate a 200-page file by position rather than reading it front to back. And when the National Archives was asked which documents in a file carry the genealogical payload, it named eight. Read those eight, then the questionnaires, then stop. Everything else is a skim.

Why the file is that long in the first place

Two things drive the size, and knowing which one applies to your file tells you what is in it.

A widow's file physically contains the soldier's file. This is the single most useful structural fact about these records and it is barely mentioned anywhere. If a veteran applied for an invalid pension before he died and his widow later claimed, the National Archives explains that his files are consolidated together with her claim into the widow's certificate file, including the earlier application files and the examining surgeons' certificates. A widow's file is not a different document from a soldier's file. It is a superset of it.

The evidentiary burden compounds. A widow had to prove a valid marriage, her husband's death, the connection between that death and his service, that she had not remarried, and frequently the identity and birth dates of minor children. Every one of those proofs generated affidavits, and any proof the Bureau doubted generated a Special Examiner's investigation on top. That is why a contested widow's claim runs to hundreds of pages while a straightforward invalid claim can be under ten.

The Archives does not publish a typical length, but it says on its own order form that pension files for the Civil War and later "can be very large and average more pages than pension files for previous wars," and it prices a full file in a block of the first hundred pages with a per-page charge after that, issuing a quote for the overflow. Nobody builds a pricing structure like that for documents that are usually short. The Archives charges considerably less for the eight-document packet than for the full file, which is worth knowing before you order, and sits alongside the other prices in what it actually costs to get an old record read.

One thing not to say, because it is not published: there is no known average. The Archives has never released a mean page count or a distribution, and neither has anyone else. Files run from under ten pages to several hundred and that is as precise as the honest answer gets.

Nobody indexed the inside of it

This is why the problem exists at all, and it is worth being exact about.

Every finding aid for Civil War pensions is a name-and-number index. The General Index to Pension Files covers 1861 to 1934 across 544 rolls, alphabetically by surname. The Organization Index runs to 765 rolls, arranged by state, then unit, then name. There is a numerical index by application and certificate number, and a separate index to remarried widows' files.

What each of those gives you is a name, a unit, a rank, an application number, a certificate number, and the state where the claim was filed. What none of them gives you is any indication of what is inside the file.

So the situation is exact: you can discover that a file exists. You cannot discover what is in it without reading it. That gap is not a gap in your research skills. It is a gap in the finding aids, and it has been there for a century.

The file is in a known order

Pension files are not shuffled. The Pension Bureau imposed a filing sequence, set out in its own Orders, Instructions, and Regulations of 1915, and for an original invalid application it ran in this order:

  1. Face brief
  2. Declarations by the invalid, in the order filed
  3. Other statements of claimants
  4. Powers of attorney
  5. Fee agreements
  6. War Department reports
  7. Other department reports
  8. Certificates of disability for discharge
  9. Evidence of prior soundness
  10. Evidence of origin
  11. Evidence of continuance, chronologically
  12. Certificates of medical examination, chronologically

Where a widow continued the claim after the soldier's death, a further block was filed after position eleven: proof of the date of his death, proof of marriage, proof of any former marriages, and proof of her widowhood or remarriage. Where minor children continued it instead, that block covers the widow's death, proof of the children's birth dates, and letters of guardianship.

Look at where the genealogy lives in that list. Positions one and two, and the widow's or minor's block after position eleven. Positions four through seven are close to pure administration and can be skipped outright. Positions ten and twelve are mixed, for reasons set out below. That means a 200-page file has a shape, and the shape tells you where to open it.

One caveat, stated honestly. The 1915 regulations describe how files were supposed to be arranged. Real files were handled, reorganised and consolidated over decades, and consolidation merged separate applications into single files. Treat the order as a strong prior about where to look first, not a guarantee about what you will find there.

The National Archives already did your triage

This is the fact that should change how you approach the file, and as far as we can tell no consumer-facing article has ever used it.

When you order from the National Archives you can buy either the full file or something called a Pension Documents Packet. The Archives describes the packet as containing, to the extent they are present in the file, eight documents that contain genealogical information about the applicant. They are:

  1. Declaration of pension
  2. Declaration of widow's pension
  3. Adjutant General statements of service
  4. Questionnaires completed by applicants
  5. "Pension Dropped" cards
  6. Marriage certificates
  7. Death certificates
  8. Discharge certificate

Consider what that list is. The institution that has custody of these files, on the order of two million of them by its own separate counts of soldiers' and dependents' cases, and has handled them for a century, was asked to pick the pages that carry the genealogy out of a file of a hundred pages or more. It picked eight document types. It charges substantially less for them than for the whole file, which is a fairly direct statement of what it thinks the remaining ninety-odd pages are worth to someone researching a family.

If you have the full file already, that list is your reading order. You do not need to buy anything to use it.

The inclusion of the "Pension Dropped" card is the interesting choice, and it is diagnostic of how the Archives thinks. It is an administrative card recording why a pension stopped, which for a widow usually means she died or remarried. A life event, dated, on a form nobody would think to look for.

What we would rank, and why

The eight-document list and the filing order are the Archives'. The ranking below is ours, because nobody publishes one, and you should treat it as our judgement rather than an institutional position.

Read closely.

The 1898 and 1915 questionnaires are the densest genealogical documents in the entire file, and they are usually one or two pages each. The Bureau circulated them to surviving pensioners. The 1898 form asked whether the pensioner was married, his wife's full name and maiden name, when and where and by whom they were married, what record of the marriage existed, whether he had been married before and to whom, and the names and dates of birth of all living children. The 1915 form went further, adding date and place of birth, current marital status, and a complete list of children living or dead with their dates of birth.

Read that last clause again. A federal form, answered in your ancestor's own hand, listing children who died before any census recorded them. There is very little else in American genealogy that does that. If the file is that of an immigrant veteran, the 1915 form's place-of-birth question is also one of the better places a European town name turns up, which we go through in your ancestor's records just say Germany.

The declarations, at positions one and two, give the veteran's name, age, residence, unit, enrolment and discharge dates, a physical description, and the nature of the injury or disease. The widow's declaration gives her residence, her maiden name, and the date and place of marriage. For many women that maiden name appears in no other federal record.

Marriage proofs, including ministers' returns and, where no civil record existed, sworn affidavits from people who witnessed the ceremony, which name those witnesses and often state how they were related.

Family Bible pages. These were physically torn from the Bible and filed as evidence of marriage or births. The Bible itself is very often long gone. The page survives in a federal file in Washington. This is the highest-value thing that can turn up and it is completely unpredictable.

Death and burial proofs, including undertakers' bills, which name the attending physician, the place of death and the burial location.

Skim, and read closely only on a signal.

Affidavits from neighbours and comrades. Most are formulaic statements about disability and are worth nothing but the affiant's name and residence. A minority are extraordinary. The signal to watch for is a claim of long acquaintance: a comrade swearing he has known the veteran since boyhood is incidentally documenting a pre-war residence and a social circle, and a neighbour proving a widow did not remarry is documenting her household for decades. Skim every affidavit for names and dates. Stop and read whenever someone claims to have known the family before the war. The affiants are a ready-made cluster of the family's friends, associates and neighbours, which is the raw material of FAN club research.

Special Examiner depositions. Present only where the Bureau doubted something, and when present they are the best genealogy in the file: sworn, transcribed, question-and-answer interviews with the claimant, her neighbours and her relatives. A contested claim is a gift. You can spot one by the presence of Bureau internal correspondence about discrepancies or conflicts.

Surgeon's certificates. Mostly filler, and there can be a great many of them in a consolidated widow's file because the whole run from the husband's invalid claim is included. But each one records a physical description, an age and a residence on a dated occasion, so the series tracks where a man lived over years. Read the first and the last. Skim the rest.

Do not read unless something specific sends you there. Powers of attorney, fee agreements with pension attorneys, War Department transmittal reports, routing slips, and repeated certificates of continuance. The face brief is worth one look, not as evidence but as a map of the file.

The same shape applies to a probate packet

If you have ever ordered an estate file from a county courthouse you have met this problem in a different costume, and the fix is the same.

The will book is not the packet. The will book holds a clerk's copy of the will and nothing else. The packet holds the entire proceeding: the petition, publication notices, letters testamentary or of administration, bonds, the inventory and appraisal, bills of sale from estate auctions, annual accountings, receipts, the final settlement, guardianship records for minor children, and any lawsuit the family filed against itself. Everything generated after the will exists only in the packet, and packets sometimes contain wills that were never copied into the book at all.

Two documents carry most of the payload. In an intestate estate, the distribution is the richest thing in American probate, because the court was legally obliged to identify every heir and their share, which means it reconstructed the family for you. And bonds are worth more than they look, because bondsmen were usually kin, and reading who stood surety for whom is a way of mapping a family that never appears in a census.

Probate has the same indexing problem as pensions, for the same reason. Will books are usually indexed, but FamilySearch notes that some probate indexes carry only the name of the deceased and not the beneficiaries inside. Original packets typically have no comprehensive index at all and are reached by decedent name or case number. A will naming your ancestor as a legatee is invisible unless you already knew whose estate to open.

Do not put a page count on a probate packet. Nobody publishes one and the range is enormous. What drives the size is knowable, though: whether the estate was contested, whether there were minors under guardianship, whether real property was sold, and how many years the administration ran. A guardianship alone can generate annual accountings for a decade.

Where the search problem has actually been solved

It would be dishonest to write this without saying that a large part of this problem got dramatically better very recently, and the pages currently ranking for it have not noticed.

FamilySearch Full-Text Search left FamilySearch Labs on 30 August 2025 and, on FamilySearch's own figure at that time, searches nearly two billion images using AI-generated transcripts. When it launched in Labs at RootsTech in 2024 it covered about a hundred million images. That is roughly twentyfold growth in eighteen months, and review pages written at launch, still ranking today, quote the hundred million figure.

What it changed is real and specific. It searches complete transcripts rather than a handful of indexed fields, which means a name is findable anywhere in a document. Deeds have always been indexed by grantor and grantee only, so every other person named in the instrument, the witnesses, the adjoining landowners, the wife releasing her dower right, was effectively unfindable. Now they are findable. For US land and deed records, and a substantial body of probate, the search problem is largely solved. Say so plainly.

Three things it does not solve.

It cannot search images that were never scanned. Civil War pension files are not in it. Fold3 put its Civil War widows' pensions collection at 22 per cent complete when we checked, and FamilySearch's wiki recorded that as of December 2011 only 3 per cent of pension records were available online. Those two figures measure different things and should not be subtracted from one another, but the direction and the pace are not in dispute, and the far larger body of soldiers' and rejected applications is not covered by that digitisation project at all. The bottleneck there is digitisation, not search.

It is full text, not indexed search. It matches the words in a transcript rather than structured fields, so you cannot ask it for a man born around 1840 in Ohio the way you can ask an indexed collection. And FamilySearch says plainly that because the transcripts are AI-generated, you may see transcription errors. Discoverability was prioritised over transcription accuracy, which is the right trade for finding things and the wrong one for citing them without looking at the image.

It does not read the file you already have. This is the important one. Full-text search answers "does this name appear somewhere in this corpus." It does not answer "I have this document, what in it matters." A found file is still an unread file, and no amount of indexing progress changes that. Recording what you decided not to read is worth a minute of your time, and it belongs with your negative findings, because future you cannot otherwise tell a page you assessed and dismissed from a page you never opened.

The arithmetic, if you were planning to just read it

The published figures for how long a human takes to transcribe a page of historical handwriting come from adjacent fields, because nobody has published one for pension files. They agree more than you would like.

The Transcribe Bentham project, analysed by Causer, Grint, Sichani and Terras in Digital Scholarship in the Humanities in 2018 across 4,364 approved transcripts, needed a figure to cost the work against paid staff and assumed an average of 45 minutes to transcribe one manuscript at £18.35 per transcript. That is a modelling assumption rather than a measurement, and the paper says so. From a different direction, Humphries and colleagues, in the study published in Historical Methods in 2025, report that student research assistants usually transcribe around five to seven pages a day.

Take the lower of those. At 45 minutes a page, a 187-page file is about 140 hours. At five hours a week that is more than half a year, for one file, for one ancestor, and the Bentham figure was chosen to be optimistic.

There is a second Bentham finding that matters more than the rate. When the project measured how long staff took to check volunteer transcripts, the mean was around three and a half minutes, but the distribution was violently uneven: in the first period, the 17 per cent of transcripts needing ten minutes or more consumed 45 per cent of all checking time. A small minority of pages eats most of the effort. Which is the whole argument for triage, from the best-documented transcription project there is.

What KleioBase does with a file like this, and what it does not

We are a tool in this space, so read this as disclosure.

A pension file usually arrives as a PDF. You can add a PDF of up to 50 MB and 100 pages on any plan, including the free one, and pick which pages become records from a page picker, so a 187-page file becomes the fifteen pages you decided to process rather than all of them. The PDF itself stays in your browser and is never uploaded to us; only the pages you select are sent. Adding pages is free, and processing is what consumes credits. Before processing you can add context, meaning the place and approximate year, and draw region boxes to point at one entry on a page. Bulk upload of a zip of images is available on the Archivist and Professional plans.

What happens after the reading is the part that matters for a file like this. A pension file names dozens of people: the veteran, his wife, her previous husband, their children, two comrades, three neighbours, a physician, an undertaker, a pension attorney. Confirming a record creates or updates a profile for each person you keep, links the record to them with their role, records the family relationships, and starts duplicate matching in the background, so the neighbour who signed an affidavit in 1887 connects to the same man in an 1880 census page you processed last year. This is the layer we argue is missing from most research workflows, set out in research logs track where you looked, nothing tracks what you found. Witnesses and associates are kept out of your main connection count but stay findable. The mechanics are in uploading records and finding duplicates.

Now the honest limits.

We cannot tell you which fifteen pages to pick. That judgement is the subject of this article and it is yours. What we can do is make processing the fifteen cheap enough that picking well matters more than picking fast.

The review step is the real cost at volume. Every processed record enters a Review state where you check the fields against the image before confirming. On fifteen pages that is manageable. On a hundred and eighty-seven it is not, which is another argument for the triage rather than against the tool.

We are not a storage service. Keep your own copy of the file. We hold the records you process, not your archive of originals.

Where to start on Saturday

Do not open page one.

Spend twenty minutes building a one-line inventory: page number, document type, date, names appearing. You do not need to read the handwriting properly to do this, only to recognise a form when you see one. Then pull the eight document types the Archives named, plus any questionnaire, and read those properly. Then skim the affidavits for anyone claiming to have known the family before the war.

You will be finished by lunch, and you will have got more out of the file than a month of reading it in order.

The Bureau clerks who assembled these files were not writing for you. They were building an evidentiary case for a payment, in a fixed order, and the family history is a by-product they filed between the fee agreement and the surgeon's report. Knowing where they put it is most of the work.

Follow KleioBase on Google

Add us as a preferred source to see our research guides more often in Google Search and AI results.

Add as preferred source

Start building your family history

Upload a record and let KleioBase transcribe, translate, and connect it - all in one place, with a research partner that remembers everything you find.

Get started

We use cookies and similar technologies. Essential cookies keep the site working and secure. We only load analytics (PostHog, including masked session replays) and marketing (Meta Pixel) with your consent. See our Privacy Policy.