# Research Logs Track Where You Looked. Nothing Tracks What You Found.

Your genealogy workflow stores two things: the searches you ran and the conclusions you reached. Everything in between, which is all of the actual evidence, sits in files you will not open again. Here is what that costs, and how to tell whether it is your real bottleneck.


In March you found an 1878 marriage record. You transcribed it, confirmed the bride and groom, and added them to your tree. The record also named two witnesses, the bride's father, and a priest.

In August you find a different record, in a village forty kilometres away, with one of those witness surnames on it.

You will not make the connection. Not because you are careless, but because the March record's witnesses exist in a scan and a transcription file that you have not opened since March, and your tree only ever knew about the bride and groom. The information is on your own hard drive. It is not in your system.

This is the most common failure in genealogy research and it has nothing to do with archives, handwriting, or access. It is a workflow problem, and it is worth being precise about its shape.

## The two things your system stores

Nearly every genealogy setup, from a shoebox to a professional's Airtable base, stores two categories of thing.

**Searches.** Where you looked, when, in what collection, and whether you found anything. This is the research log, and the entire established genre of genealogy organization is about it: the printable templates from Family Tree Magazine and Cyndi's List, the Airtable systems professionals build, and purpose-built tools like Goldie May, which logs the pages you visit into a project and has since added AI that reviews an objective for gaps and suggests collections you have not searched. The purpose of the layer is to stop you repeating a search you already ran and to record the searches that came back empty.

**Conclusions.** Who you have decided is related to whom. This is the tree, and it is what Ancestry, MyHeritage, FamilySearch, RootsMagic and every GEDCOM file hold. A tree is an assertion set. It says this man is that woman's father, and it hangs sources off that claim as support.

Both are worth keeping. Neither is what is missing.

| Layer | What it holds | The question it answers | Where researchers keep it |
| --- | --- | --- | --- |
| Searches | Collections consulted, dates, results including the empty ones | Have I already looked here? | Research log: printable forms, spreadsheets, Airtable, browser extensions |
| **Evidence** | **Every person, date, place, age and stated relationship written in every record already read** | **Where have I seen this name before?** | **Mostly nowhere** |
| Conclusions | Asserted relationships, with sources attached as support | Who is related to whom? | The tree: Ancestry, MyHeritage, FamilySearch, RootsMagic, GEDCOM |

## The layer nobody keeps

Between the search and the conclusion sits the evidence: every name, date, place, occupation, age and stated relationship written in every record you have already read.

That layer is enormous and almost nobody holds it in queryable form.

Consider one ordinary Catholic baptismal entry. It names the child, the father, the mother, the mother's maiden name, two godparents, sometimes a godparent's spouse, the officiating priest, and a house number. That is seven or eight people and a location, in four lines of Latin. Of those, most researchers record one, occasionally three.

The other five are not lost, exactly. They are in the scan. They are possibly in a transcription. But they are not in anything that will raise its hand in six months when a name comes back around, which means that functionally they are lost, and you will pay to find them again.

Multiply by every record you have processed. A researcher three years into a project has typically read several hundred records and entered a fraction of the people named in them. The gap between what you have read and what your system knows is usually wider than the gap between what your system knows and what you have not yet found.

## Why the gap exists

It exists because it used to be rational.

When reading a record was the expensive step, being selective about what you extracted was correct. If a page takes an hour to transcribe and another twenty minutes to enter, you extract the two people you came for and you move on. Recording eight people per record, with their stated ages and their relationships to each other, would have tripled the work for people you did not care about yet.

That calculation has changed. Transcription is now the cheap part of the process, and it got cheap fast. FamilySearch's full-text search went from about a hundred million images at launch in 2024 to 1.2 billion two years later, drawn from 29 countries rather than the original three. Ancestry shipped document transcription in April 2025, MyHeritage shipped Scribe in March 2026, and Transkribus has been doing this for institutions for years. Whatever else you think about those tools, the direction is not ambiguous: reading a record is no longer the constraint.

The workflow did not update. Most researchers now transcribe faster and still extract selectively, into the same tree they were keeping before, which means the transcription capacity mostly goes into producing more files they will not re-read.

## The cost that does not show up

Here is the part that makes this urgent rather than merely untidy.

Reading records is linear work. Ten records take ten times as long as one. If you double your reading speed, you halve that cost, and the relationship holds forever.

Reconciling records is not linear. Every new person you add has to be considered against everyone you already have, because the question "have I seen this person before" is asked against the whole corpus. Two hundred people in your files means roughly twenty thousand possible pairs. Four hundred people means roughly eighty thousand. You doubled the corpus and quadrupled the checking.

You do not feel this as a wall. You feel it as research that gets slower the better it goes, which is easy to misread as the records getting harder. Sometimes they are. Often what has actually happened is that your corpus outgrew your memory, and memory was doing the reconciliation.

The consequence is uncomfortable and worth stating plainly: **making transcription faster makes this worse.** More records read, at the same extraction rate, produces more unreconciled material. The linear step got cheaper and the quadratic step got bigger.

## Which direction the checking runs

There is plenty of automation aimed at this, and it all runs the same way round.

| System | Starts from | Checked against | Direction |
| --- | --- | --- | --- |
| Ancestry Hints | A person in your tree | Ancestry's licensed collections | Outward |
| MyHeritage Record Matches and Smart Matches | A person in your tree | MyHeritage's records and other users' trees | Outward |
| FamilySearch Research Helps | A profile in the shared tree | FamilySearch's records and the tree itself | Outward |
| Research assistants and logs | A research objective or a search you ran | Catalogues of collections you have not searched | Outward |
| **KleioBase** | **A record you just read** | **Every record you have already read** | **Inward** |

Disclosure, since it is our table: we make the last row. KleioBase is a genealogy research platform that transcribes handwritten records in any language, turns every person named in them into a profile, and checks each new record against everything you have already read. We went looking for another product to put in that row and did not find one, which is the whole reason this article exists.

Every row but the last starts from a conclusion you have already committed to and looks out at material you have not seen. That is a useful thing to automate and the platforms are good at it. It is also the easier problem, because the corpus on the other side is indexed, owned by one company, and identical for every user.

The last row is the one that compounds, and it is the one almost nothing runs. It starts from the record in front of you and checks it inward, against your own accumulated material, including every person you transcribed and never promoted into your tree. The witness. The godparent's spouse. The neighbour who signed. The corpus on the other side of that comparison is different for every researcher, which is exactly why no record platform has an incentive to build it and exactly why it is worth more to you than a hint is.

These are not substitutes. Outward checking tells you where to look next. Inward checking tells you that you already looked, five months ago, and did not notice.

## The methodology everyone recommends and almost nobody runs

There is a well-established answer to hard research problems, and it has been the standard advice for decades. Elizabeth Shown Mills named it the FAN principle: friends, associates and neighbours. When you cannot trace an individual directly, you trace the cluster around them, because people moved, married, witnessed and settled in groups.

Ask any experienced genealogist how to break a brick wall and cluster research will be in the first three answers. Now ask how many of them systematically record every witness, godparent and informant they encounter.

The method is not unpopular. It is unaffordable. Cluster research requires exactly the layer described above, held for years before it pays off, and building it by hand costs more than most brick walls are worth.

That is a clerical constraint, not an intellectual one. It is the kind of constraint that software removes.

## A diagnostic

Before concluding that this is your problem, test it. Five questions, answered honestly about your own current setup.

1. **Can you list every record you hold that mentions a given surname, in under a minute, including records where that surname belongs to a witness rather than a principal?**
2. **Do you know how many distinct people are named in your files but do not exist in your tree?** Not the number. Whether you could find out at all.
3. **When you open a new record, what do you check it against?** Your files, or your memory of your files?
4. **How often do you re-open a record you have already read, to remind yourself what was in it?**
5. **Can you state what you have ruled out?** Not what you have not found. What you have looked for, properly, and established is not there. That distinction is [negative evidence](/blog/negative-evidence-in-genealogy), and it is only usable if it is recorded.

If questions 1 and 2 are answerable, your bottleneck is genuinely finding and reading records, and you should spend your money on access and your time on paleography.

If they are not, then more records will not help you as much as you expect, because you are not using the ones you have.

## What an evidence layer actually has to do

Suppose you decided to fix this. Tool-agnostic, these are the requirements. They are worth holding onto because they are also the right questions to ask of any product that claims to solve it, including ours.

**Every person mentioned becomes an entity, not a note.** The witness on the 1878 marriage has to be a thing the system knows about, with a name, an age if stated, a place, and a link back to the record and the image. A transcription that sits in a text file is not this. Neither is a scanned PDF with searchable text, because search finds strings, not people.

**Reconciliation happens on ingest, and it runs inward.** Per the direction table above, the check has to fire when the record enters, without you deciding to run it, and it has to run against everything you have already read. If it requires you to remember to look, it has the same failure mode as memory, which is the thing it was supposed to replace.

**Uncertainty survives.** Two records giving different ages for the same man is information. A system that silently picks one, or averages them, has destroyed the evidence and handed you a conclusion. The contradiction should be visible and should stay visible until you resolve it deliberately.

**Extractions stay attached to their images.** An AI reading of a handwritten page is a reading, not a fact. It can be wrong about a name, and names are exactly where it is most likely to be wrong, because unlike the formulaic parts of a record a proper noun has nothing around it to constrain the guess. If you cannot get from a stated fact back to the pixels in one step, you cannot cite it and you should not trust it.

**The system can say what is missing.** Given a person's known dates, place and jurisdiction, the records that should exist are largely computable. A man who married in Kovno gubernia in 1878 and died in 1911 should have a death record, and the archive that would hold it is knowable. This is deterministic work rather than a matter of AI judgment, and it turns an open-ended question into a list. Several tools now attempt some version of this, with varying scope.

**The tree becomes an output.** If the evidence layer is real, the tree stops being the thing you type into and becomes a view of what the evidence supports. That is the point of the whole exercise.

**Your data comes out in a standard format.** Whatever the system holds, you should be able to leave with it. GEDCOM is the only format the field agrees on, so import and export in it is the minimum bar.

### What that looks like on the 1878 record

Go back to the marriage from the first paragraph.

With an evidence layer, processing it produces seven profiles rather than two. The bride and groom, the bride's father, both witnesses, the priest, and the groom's father if he is named. Each one carries the record, the image, and the age or occupation the record states, and none of them are asserted to be related to anyone except as the document itself says.

Five months later the second record arrives with a matching surname on it. The check runs on ingest. What comes back is not an answer. It is a flag: this surname, in this district, appears on a record you processed in March, as a witness, aged 34. That is a lead you would not otherwise have had, and it costs nothing to produce because the work was done in March.

Nothing here is clever. It is bookkeeping. The reason it does not happen is that the bookkeeping is too expensive to do by hand, which is a statement about clerical cost rather than about research skill.

## Where this approach goes wrong

Four failure modes, stated up front, because a page that describes only the upside is an advertisement.

**Automated matching produces false positives, and a bad queue is worse than no queue.** Match confidence is not uniform. High-confidence candidates are usually right; low-confidence candidates, particularly ones based on name similarity alone, are usually wrong. A system that shows you everything it can think of will burn more of your time than it saves, and it will burn your trust first. Judge any matching feature on what it declines to show you.

**Suggestion engines overproduce.** The same is true of gap detection and research hints. It is trivial to generate a thousand things you could look at. The hard problem is ranking, and any tool in this space, ours included, is more likely to be bad at ranking than bad at generating.

**It does not do the adjudicating.** When two records conflict, deciding which is right is analysis, and analysis is yours. The Genealogical Proof Standard has not changed, and the Board for Certification of Genealogists' 2024 interpretation is explicit that AI may not transcribe, translate, abstract or analyse the supplied document in certification portfolio work. A system that surfaces a conflict has done its job. A system that resolves one quietly has done damage.

**Garbage in scales too.** Automatic cross-referencing across a corpus with a bad transcription in it will propagate that error further and faster than a manual workflow would. The image next to the transcription is not a nicety.

## What we built

KleioBase is this shape, and this article is the reasoning behind it rather than an afterthought.

Records go in through the [upload workspace](/docs/uploading-records), in whatever language and script they are written in. Every person named in a record becomes a profile in a [knowledge base](/docs/knowledge-base), witnesses and godparents included, each one carrying the record and image it came from. Every new upload is checked inward against everything already there, across name spellings and languages, and the candidates come back scored rather than merged. That last row of the direction table is the whole reason the product exists. Conflicts between records are flagged instead of resolved. An Expected Records engine computes which documents should exist for a person given their dates and jurisdiction, and the [Research Companion](/docs/research-companion) answers questions against the knowledge base with citations back to the records. The [family tree](/docs/family-tree) is drawn from what is in the knowledge base rather than typed in. GEDCOM goes [both ways](/docs/importing-and-exporting).

What it is not: it is not a record database, it does not hold anyone's archive, and it does not decide what is true. It reads what you give it, keeps all of it, and tells you when something you already have is relevant to something you just found.

## Where to start, whatever you use

None of this requires our software or anyone's.

Start by changing one habit: when you process a record, record every person named in it, not only the ones you came for. Witnesses, godparents, informants, the neighbour who signed. Give each one the record they came from. This is more work per record and it is the only version of this that compounds.

Then make your evidence queryable by person rather than by file. A folder of PDFs named after their collection is a filing system. A table where one row is one person-mention is a research instrument, and you can build a usable version of it in a spreadsheet in an afternoon.

Then record what you ruled out, with enough specificity that future-you can tell a search that failed from a search that was never run.

Do those three things and the tooling question becomes secondary, which is the correct order to take it in. The reason to automate this is not that automation is interesting. It is that the manual version is expensive enough that almost nobody sustains it, and the people who do sustain it are the ones whose research does not stall.

If you want the shorter version of what changes when the evidence layer exists, our post on [the family tree that shows what is missing](/blog/family-tree-that-shows-what-is-missing) is the same argument told through one feature. For where AI actually helps and where it does not, see [how AI is transforming genealogical research](/blog/how-ai-is-transforming-genealogy).

Canonical: https://kleiobase.com/blog/genealogy-research-evidence-layer
