Test it before you trust it.
Evaluating an AI assistant in six lessons. A course book and six notebooks that run offline, for the person who has to say whether an assistant is good enough to launch.
What it is
A way to test an AI assistant that answers questions from documents, before you decide to launch it.
- A course book of 35 pages, as a PDF.
- Six Jupyter notebooks, one for each lesson. They run with no network and no API key.
- A small teaching dataset, recorded in advance, that comes with the course.
- Three exercises at the end of each lesson, with worked answers in a separate notebook.
- A one-page checklist and a page of sources.
The dataset is made up for teaching. The library in it does not exist, and its numbers say nothing about any real product.
Not for sale yet.
Read lesson 1, the free sampleWho it is for
An analyst or a team lead who can read Python, has an AI assistant from any vendor answering questions from documents, and must decide whether it is good enough to launch.
You do not need an account with any AI vendor to take the course. You do not need to be a statistician.
The six lessons
- Write the questions before you look at the answers. Build a test set and hold part of it out.
- Grading answers. Exact match, a written rubric, and why two graders must agree.
- Is the difference real? Intervals and paired comparisons with small samples, with the sample size said out loud.
- Grounded or not. Check each answer against its source passage, and count the claims the passage does not support.
- Failure hunting. Slice the results, plant a canary that must fail, and see why a check that cannot fail proves nothing.
- The launch memo. One page that states what was measured, on how many held-out questions, what was not measured, and the decision.
Each lesson is a chapter in the book and one notebook. The course follows three lines from How we work:
- Ground every answer in its source.
- Measure on held-out sets before launch.
- Hand over code, standards and training.
What is in the box
- The course book: 35 pages, with a contents page, ten figures, a one-page checklist and a sources page.
- Six lesson notebooks and one notebook of worked answers to all eighteen exercises.
- The teaching dataset: 8 documents in 40 passages, 150 test questions, 300 recorded answers from two made-up assistant versions, two graders' marks on every answer, 346 labeled claims and 12 planted bad answers.
- The script that made the dataset, and a page that describes every column.
- A README that says how to run everything.
To run the notebooks you need Python 3.10 or later with numpy, pandas, matplotlib and Jupyter. Nothing in the course calls an assistant.
What it is not
- It is not a statistics textbook. It uses four tools and explains each only as far as you need to use it.
- It is not about any vendor's product, and it does not tell you how to build or tune an assistant.
- It does not test your assistant for you. It teaches a method on made-up data, and you apply it to your own.
- It does not tell you what any law, contract or regulator requires of you.
- It is a book and notebooks that you work through on your own. There are no live sessions.
Lesson 1, in full
This is the free sample: lesson 1 as a PDF excerpt of the course book, and its notebook with the three data files it reads. The notebook runs offline.
The PDF has 7 pages. The notebook comes as a zip file with its data and a README. The worked answers to the exercises are in the full course, not in the sample.