The AIM Framework: How Family Historians Stay in Charge of AI Research

AIM is the Chronicle Makers method for working with AI on family history research: Assess, Interact, Measure. It replaces "trust but verify."

Share
Three woodcut letter blocks of the letters AIM with a wood chisel and shavings in front of it.

The AIM Framework is the Chronicle Makers method for working with AI on family history research: Assess, Interact, Measure. You set the terms before the tool runs, you hand it your records rather than open questions, and you measure everything that comes back against sources and a method you stated in advance.

AIM replaces "trust but verify." It sounds wise and contains no testable instruction — trust how, checked where, against what. AIM is the answer.

This page is the framework's permanent home. It covers what each part means, where these AI tools commonly fail, and how to know whether your own checking is working.

Fluent output and correct output are unrelated

Writing quality carries no signal about the research quality. I continually receive stories from family historians where Claude, ChatGPT, or Gemini told them facts and connections that sounded correct in the chat, but when checked, fell apart.

A language model produces prose of consistent quality regardless of whether the underlying evidence supports the claim. A weak identification and an airtight one come back in the same tone, with the same confident sentences. A result that sounds finished is the single most common way a family historian accepts something wrong.

Everything in AIM follows from that. The framework puts the family historian in charge at three points: before the tool runs, while it works, and after it answers.

Assess: the method you choose before anything runs

Assess comes first. What is the task or job I am giving AI to do, and can AI do it?

Ask a model to reconstruct families from two thousand church entries and it will not ask you how. It will decide how: what counts as a match, when two entries describe the same household, what to do with a name spelled three ways. Then it will explain its approach afterward if you ask. Reading the model's stated approach and agreeing with it is not the same as having chosen the approach.

The way a task or job gets completed is a decision, and it belongs to the family historian. It gets stated before the work starts, pointed at the standard it satisfies, and written down where it will be applied. What comes back is then measured against the stated method — did it do what I said, on every record — rather than against a rationale the tool supplied for itself.

The question: did I give the method, or did I approve one?

Interact: hand it records, not open questions

The second part of AIM is about what you give the tool, because that is what decides whether the answer can be trusted at all.

These tools are strong when working with material you hand over. Handwriting transcription is now genuinely good on eighteenth and nineteenth century American records, and it is fast. Translation of record-book Latin, German and Swedish is strong. Extracting names, dates and places into a table from a batch of images works consistently. Historical context — what a county looked like in 1798, what a term in a probate inventory meant — comes back accurate far more often than not.

A transcription of a census record can be completed in minutes.

Lines 35–38 — Dwelling 317 / Family 317  Line	Name	Age	Sex	Occupation	Real estate / Personal estate	Birthplace 35	Hollman, Wilhelm A.	43	M	Farmer	~$1,800 / ~$600 (numerals hard to pin down precisely on this copy)	Penn. 36	— Margaret A.	41	F	Keeping House	—	Penn. 37	— Henry S.	12	M	At home	—	Penn. 38	— Alice M.	4	M (see note below)	At home	—	Penn.  All four are marked "W" for color. Wilhelm A. also has a mark in column 19, off to the right.
Transcription of page 41 of the 1870 US Census population schedule for Plymouth Township, Montgomery County, Pennsylvania with Claude Sonnet 5.

This transcription combined with the original census enumerator instructions, allows Claude to provide the next level of analysis for the information on the form, also in minutes.

What the form is actually telling you  Columns 1–2 (dwelling/family numbers) only get written when the enumerator starts a new house or a new family. The instructions say they're "numbered consecutively, in order as visited" — a fresh number means a fresh unit, a blank means "still counting the same one as above." That's why lines 36–38 carry no numbers: they're part of family 317, opened on line 35. The dash in front of Margaret A., Henry S., and Alice M. isn't a redaction — it's exactly what the instructions describe: "So long as the family name remains the same for the several members, it need not be repeated, provided a clear horizontal line be drawn in the place it would occupy." It just means "same surname as the line above." "Farmer" vs. "Farm Laborer" wasn't a casual choice of words. The instructions call it out specifically: "Be very particular to distinguish between farmers and farm laborers. In agricultural regions this should be one of the points to which the Assistant Marshal should especially direct his attention." Wilhelm A. is the one recorded with property values in columns 8–9; Jacob has neither. That split — owner/operator with real estate, versus laborer with none — is the distinction the Census Office was after. "Keeping House" vs. "Domestic Serv't" look similar but were meant to mean different things. Per the instructions, "'house-keeper' will be reserved for such persons as receive distinct wages... Women keeping house for their own families... will be entered as 'keeping house,'" while paid household help elsewhere is to be reported simply as "domestic servants." So Margaret A. was running her own household unpaid; Anna was doing paid domestic work — the page doesn't say for whom, just that it was a job, not her own home.
Claude Sonnet 5 analysis of information provided in fields, referencing the Instructions for US Marshalls for the Ninth Census, provided by the US Census Bureau.

What those tasks share is that the tool is reading, converting or retrieving from something real in front of it. Failure rates climb as soon as the answer has to be constructed — and they climb fastest when the request is an open ask with nothing attached. A tool holding your record can be wrong about a word. A tool holding nothing can be wrong about everything, fluently.

Measure: where AI fails, by task

The third part of AIM is human looking at the AI output and deciding — In what ways did AI correctly accomplish what I asked it to do? And it what ways did it fail?

AI does fail and it fails in specific ways, but not consistently:

Transcription fails at the illegible spot, not across the document. Where the ink is gone or visibility is impossible, the model supplies a plausible word rather than reporting that it cannot read it. The completion matches the surrounding language, so it reads as correct. Each word must be checked against the words in the original.

Extraction invents to fill the column. Give a tool a table structure and it will populate every cell. A blank in the record becomes an inference in the table, and AI will fill in data to make it look complete — a calculated birth year sitting in the same column as a written date of birth. Recorded and derived have to stay visibly separate.

Summarizing drops the exception. Models compress toward the pattern, and in family history the outlier is usually the whole point. The one entry that breaks the sequence, the child who appears once and never again, the witness with the wrong surname are all examples of what a summary can smooth away.

Correlation is agreeable. Tell a model you believe two records are the same man, and it will build you a coherent account of one life. It is working from your framing and it will not push back unprompted. Two people of the same name about the same age in the same county being combine into one individual, is a standard failure in genealogy.

Conflict detection catches the internal, misses the contextual. Two dates that disagree inside your documents will get flagged. A conflict that requires knowing the county line moved in 1800, or that the record was kept under the old calendar, will not be, unless the tool was given that context during Assess. Relying on the tool's training data for the context necessary to flag conflicts is building a house on sand.

Citations are the highest-risk output. They come back in perfect format, which is precisely the problem. Formatting correctness and existence are independent. Every citation gets measured against the source consulted, always.

Drafted proof arguments read as proof regardless. The narrative will be organized, it will move from evidence to conclusion, and it will sound like something that met a standard. Whether the evidence underneath supports it is a separate question, no matter how great it sounds.

How to measure your own checking

AI is impressive when it completes work in minutes that used to take us hours. It is presented to us perfectly formatted and look authoritative. Sometimes AI has us feel like we are a "genealogy genius" and that feeling is red flag to immediately pause and check the work.

Take material you have already worked and measured — a set of records where you know every answer. Run it through the tool as if it were new. Count what came back right, what came back wrong, and, most usefully, which of the wrong answers you would have accepted if you hadn't already known. That last number is the one that tells you where your own way of checking AI's output is lacking..

Rerun it when you change tools or move to a new record set. Model behavior shifts between releases, and an accuracy figure from last year is a number about a tool that no longer exists.

Where the AIM Framework fits

Chronicle Makers publishes each of its frameworks at a permanent address, and each one answers a different question.

The Soundness Standard governs when the research is done. The Research Quality Check runs that question against one ancestor at one moment. AIM governs how the person stays in charge while the work is happening — before, during, and after every exchange with a tool. All of them have to be satisfied: a result can be accurate and the work still unfinished, and work can feel finished while resting on a transcription nobody opened the image to check. If the question underneath all this for you is permission rather than method, that one has its own answer.

Nothing here is a reason to use these tools less. I use them daily, on my own Pennsylvania research, and the volume they make possible is the reason projects that stalled for years are finishing now. But volume without measuring produces a larger quantity of unchecked conclusions, faster, and family history does not need any more unverified ancestors and family trees.

The tools are strong enough now that the constraint has moved. It is no longer what AI can do. It is whether the person reading the output set the terms before the prompt was ever sent, along with whether they checked the output and reasoning.


The AIM Framework — a framework from Chronicle Makers · Denyse Allen · chroniclemakers.com/aim-framework/ · First published September 2026.