The AIM Framework: How Family Historians Stay in Charge of AI Research
AIM is the Chronicle Makers method for working with AI on family history research: Assess, Interact, Measure. It replaces "trust but verify."
The AIM Framework is the Chronicle Makers method for working with AI on family history research: Assess, Interact, Measure. You set the terms before the tool runs, you hand it your records rather than open questions, and you measure everything that comes back against sources and a method you stated in advance.
AIM replaces "trust but verify." It sounds wise and contains no testable instruction — trust how, checked where, against what. AIM is the answer.
This page is the framework's permanent home. It covers what each part means, where these AI tools commonly fail, and how to know whether your own checking is working.
Fluent output and correct output are unrelated
Writing quality carries no signal about the research quality. I continually receive stories from family historians where Claude, ChatGPT, or Gemini told them facts and connections that sounded correct in the chat, but when checked, fell apart.
A language model produces prose of consistent quality regardless of whether the underlying evidence supports the claim. A weak identification and an airtight one come back in the same tone, with the same confident sentences. A result that sounds finished is the single most common way a family historian accepts something wrong.
Everything in AIM follows from that. The framework puts the family historian in charge at three points: before the tool runs, while it works, and after it answers.
Assess: the method you choose before anything runs
Assess comes first. What is the task or job I am giving AI to do, and can AI do it?
Ask a model to reconstruct families from two thousand church entries and it will not ask you how. It will decide how: what counts as a match, when two entries describe the same household, what to do with a name spelled three ways. Then it will explain its approach afterward if you ask. Reading the model's stated approach and agreeing with it is not the same as having chosen the approach.
The way a task or job gets completed is a decision, and it belongs to the family historian. It gets stated before the work starts, pointed at the standard it satisfies, and written down where it will be applied. What comes back is then measured against the stated method — did it do what I said, on every record — rather than against a rationale the tool supplied for itself.
The question: did I give the method, or did I approve one?
Interact: hand it records, not open questions
The second part of AIM is about what you give the tool, because that is what decides whether the answer can be trusted at all.
These tools are strong when working with material you hand over. Handwriting transcription is now genuinely good on eighteenth and nineteenth century American records, and it is fast. Translation of record-book Latin, German and Swedish is strong. Extracting names, dates and places into a table from a batch of images works consistently. Historical context — what a county looked like in 1798, what a term in a probate inventory meant — comes back accurate far more often than not.
A transcription of a census record can be completed in minutes.

This transcription combined with the original census enumerator instructions, allows Claude to provide the next level of analysis for the information on the form, also in minutes.

What those tasks share is that the tool is reading, converting or retrieving from something real in front of it. Failure rates climb as soon as the answer has to be constructed — and they climb fastest when the request is an open ask with nothing attached. A tool holding your record can be wrong about a word. A tool holding nothing can be wrong about everything, fluently.
Measure: where AI fails, by task
The third part of AIM is human looking at the AI output and deciding — In what ways did AI correctly accomplish what I asked it to do? And it what ways did it fail?
AI does fail and it fails in specific ways, but not consistently:
Transcription fails at the illegible spot, not across the document. Where the ink is gone or visibility is impossible, the model supplies a plausible word rather than reporting that it cannot read it. The completion matches the surrounding language, so it reads as correct. Each word must be checked against the words in the original.
Extraction invents to fill the column. Give a tool a table structure and it will populate every cell. A blank in the record becomes an inference in the table, and AI will fill in data to make it look complete — a calculated birth year sitting in the same column as a written date of birth. Recorded and derived have to stay visibly separate.
Summarizing drops the exception. Models compress toward the pattern, and in family history the outlier is usually the whole point. The one entry that breaks the sequence, the child who appears once and never again, the witness with the wrong surname are all examples of what a summary can smooth away.
Correlation is agreeable. Tell a model you believe two records are the same man, and it will build you a coherent account of one life. It is working from your framing and it will not push back unprompted. Two people of the same name about the same age in the same county being combine into one individual, is a standard failure in genealogy.
Conflict detection catches the internal, misses the contextual. Two dates that disagree inside your documents will get flagged. A conflict that requires knowing the county line moved in 1800, or that the record was kept under the old calendar, will not be, unless the tool was given that context during Assess. Relying on the tool's training data for the context necessary to flag conflicts is building a house on sand.
Citations are the highest-risk output. They come back in perfect format, which is precisely the problem. Formatting correctness and existence are independent. Every citation gets measured against the source consulted, always.
Drafted proof arguments read as proof regardless. The narrative will be organized, it will move from evidence to conclusion, and it will sound like something that met a standard. Whether the evidence underneath supports it is a separate question, no matter how great it sounds.
How to measure your own checking
AI is impressive when it completes work in minutes that used to take us hours. It is presented to us perfectly formatted and look authoritative. Sometimes AI has us feel like we are a "genealogy genius" and that feeling is red flag to immediately pause and check the work.
Take material you have already worked and measured — a set of records where you know every answer. Run it through the tool as if it were new. Count what came back right, what came back wrong, and, most usefully, which of the wrong answers you would have accepted if you hadn't already known. That last number is the one that tells you where your own way of checking AI's output is lacking..
Rerun it when you change tools or move to a new record set. Model behavior shifts between releases, and an accuracy figure from last year is a number about a tool that no longer exists.
Where the AIM Framework fits
Chronicle Makers publishes each of its frameworks at a permanent address, and each one answers a different question.
The Soundness Standard governs when the research is done. The Research Quality Check runs that question against one ancestor at one moment. AIM governs how the person stays in charge while the work is happening — before, during, and after every exchange with a tool. All of them have to be satisfied: a result can be accurate and the work still unfinished, and work can feel finished while resting on a transcription nobody opened the image to check. If the question underneath all this for you is permission rather than method, that one has its own answer.
Nothing here is a reason to use these tools less. I use them daily, on my own Pennsylvania research, and the volume they make possible is the reason projects that stalled for years are finishing now. But volume without measuring produces a larger quantity of unchecked conclusions, faster, and family history does not need any more unverified ancestors and family trees.
The tools are strong enough now that the constraint has moved. It is no longer what AI can do. It is whether the person reading the output set the terms before the prompt was ever sent, along with whether they checked the output and reasoning.
The AIM Framework — a framework from Chronicle Makers · Denyse Allen · chroniclemakers.com/aim-framework/ · First published September 2026.