User Research

Who was spoken to, what they said, and what changed in the build because of it. Nobody, nothing, and nothing.

No user has been interviewed, and no user feedback has changed this product. Outreach has gone to named regulatory-affairs people on live dockets and has produced no replies and no calls. The interview table below is empty, and it stays empty until somebody talks to us.

What was done instead is on this page, labelled as a substitute every time it appears. A synthetic user was built from 102 real public filings and interviewed by the person who built the product. That is a rehearsal, not evidence. It generates questions to ask; it cannot answer them.

Where the numbers live. Attempts, addresses and their provenance are logged in docs/mrd.html; what a conversation actually said would live here. The split stops the two drifting into different versions of one story.

What actually happened

Measure2026-08-04
Messages sent45. Seven on 2026-08-03, then 38 through 2026-08-04, paced at five to thirteen minutes with one recipient per employer first.
Replies0.
Interviews held0.
Calls booked0.
Bounces4, all confirmed by the mailer daemon.

Where the 45 addresses came from

Every address was read off a certificate of service in a real docket. No list was bought, no address was scraped from a company website, and the three that were pattern guesses are marked as guesses in ADR-63. The population is checkable from this repository rather than from anybody's memory: data/real/ holds 102 filings across 19 dockets and eight state commissions — Georgia, Indiana, Kentucky, Missouri, North Carolina, Ohio, Utah and Virginia — naming 48 distinct filing organisations.

Kind of organisation on those filingsCount
A utility or its outside counsel26
Commission staff14
Consumer advocate or state consumer office5
Large-load customer intervening on rates2
Other intervenor or advocacy group1

This is the population the addresses were drawn from, derived by counting the provenance records in data/real/. It is not a breakdown of the 45 people who were written to, by title or by seniority: that list lived in a mail client and never in this repository, so it cannot be reconstructed here without writing from memory. The addresses came from organisations that file in rate and interconnection proceedings, which is the population this product is built for, and no name, address or private detail of any recipient appears anywhere in this repository.

Targeting was not the problem. These were the right organisations, reached at addresses they had themselves put on a public filing. Forty-five well-aimed messages produced nothing, which says the channel was wrong rather than the aim — cold email to people who file for a living, from an address they have never seen, inside a two-day window. An introduction through somebody they already know is worth more than another forty-five.

Why nobody replied, as a guess. Cold outbound from a personal address to people who file for a living, on a two-day window, over a weekday. Regulatory affairs is not a function with spare afternoons. No warm introduction was available and none was tried, because there was nobody to ask. The sample is too small and too fast to conclude anything about the outreach itself: 45 messages with zero replies inside 36 hours is consistent with a bad pitch, a bad list, a bad channel or an ordinary response lag, and nothing here distinguishes them.

What the outreach taught, which is not the same as what a user taught

Four addresses bounced, and all four carried the label "filing-verified"

Every address was read off a certificate of service — a document the parties file themselves, under oath, so that others can reach them about that docket. That is the best source available and better than any contact broker. Four of them bounced anyway.

The lesson, as a rule. A service list records who was reachable when it was filed, not who is reachable now. "Verified" is a timestamp against a version, not a permanent property of a fact. The clearest case is in the record twice over: a July 2021 certificate of service in IURC Cause No. 45591 gives AES Indiana's counsel as tnyhart@btlaw.com; the April 2026 filing in Cause No. 46394 gives tnyhart@taftlaw.com. Same person, same role, different firm. Both were honestly verified against a public filing and only one works. Nothing announced the change.

This changed the product, and it must not be reported as user feedback. It is written up as principle 28 in docs/best-practices.html and as ADR-27 in docs/.ai/decisions.html, and it is the argument for scoping every citation to a proceeding version rather than to a proceeding. It also produced a real code change: verify_citation_for_version never consulted DocumentVersion.source_sha256, so its verdict was bound to a version id rather than to that version's bytes, and an out-of-band edit far from the cited offsets left a citation verifying clean. It refuses on mismatch now. That is a lesson from the method, not from a person. No regulatory-affairs person said it. Putting it in the "what users told us" column would be the same defect this product exists to catch — a confident claim with the wrong evidence behind it.

What it is legitimately for. Three things and no more. Rehearsing questions, so a real twenty minutes is not spent discovering the obvious. Generating candidate pains to go and check against a real source. Answering a design question when no real user is available, with the guess marked as a guess in the ADR that relies on it.

What it cannot do.

It cannotWhy not
Settle who the user isThe corpus is a record of who files. The user this product is for mostly reads, and often never appears on a service list. An earlier pass over 39 outreach targets found only 13 utility-side and matching the profile; the rest were commission staff, consumer advocates, outside counsel and consultants. Docket mining is biased against ever answering H1.
Price anythingNothing in a public filing says what a seat costs, what a budget line looks like, who signs the order, or whether this is bought by the department or by IT. A number invented here would land in the MRD and be indefensible.
Say what she would buy, adopt or abandonWillingness to pay, switching cost and what tool already sits in the workflow are all unobservable from the record. This bears directly on H2: if she already has a redlining tool she trusts, the corpus would not show it.
Give a frequency or a base rateThe corpus was assembled by hunting for errata and version pairs. Every rate derived from it is inflated by construction. "How often does this bite you" is a question for a person.
See anything that never produced a filingThe public record is the tail end of the work. The meeting where materiality was argued, the mail that routed it, the deadline nearly missed — all invisible. If the largest pain in the job never generates a document, this method cannot find it.
Disagree usefullyThe failure that matters most. The persona was built by the same mind that built Verbatim, from a corpus that mind selected. It produces pains Verbatim already claims to solve, because that is what it was assembled from. Treat its agreement as worth nothing and its disagreement as mildly interesting.

The three standing rules on its output. Nothing it says moves into the interview table below. Where it informs a decision, the ADR says the input was synthetic — "the synthetic persona flagged X, which we then checked against [filing]" is legitimate; "users told us X" is not. Every claim it makes is a line on the interview script, not a line in the PRD.

The one genuinely useful thing the rehearsal produced. The synthetic interview undercut the product seven times, and one undercut is worth carrying into a real conversation: that a filer's own redline already gets most of the job done, so the wedge is not producing a redline but ranking one. A 378-page Kentucky reissue carried roughly 1,370 lines of re-layout churn around a single substantive change, and a Utah redline invented three changes that never happened while silently passing over one that did. Those are facts from filings, checkable in data/real/. Whether they are the pain a real analyst feels is precisely what nobody has been asked.

The substitute, and what it is not

docs/synthetic-user.html and docs/synthetic-interview.html stand in for interviews that did not happen, and must be named as substitutes wherever they are cited. Nobody said any of it. The persona and her employer are invented. What is not invented is the evidence under each trait: every one traces to a line in data/real/ — 102 public filings, roughly 4,023 pages, across eight commissions, each with a provenance file naming its docket, filer, date, source URL and SHA-256. Traits with no source do not appear, and traits that are guesses are tagged as guesses.

Interviews

One block per conversation, filled the same day, quotes verbatim. The table is empty.

#Who / role / sectorDateMediumOutcome
1
2
3

What changed because of user feedback

What we heardWho said itWhat we changed, or why we did not
Nothing. No user has said anything, and inventing a row would cost more than the empty table does.

What the outreach method changed, kept separate on purpose. Two changes below are real and neither is user feedback.

What happenedWhat changed in the buildWhy this is not user feedback
Four filing-verified addresses bouncedPrinciple 28 and ADR-27: a verified fact has a shelf life. verify_citation_for_version now binds its verdict to DocumentVersion.source_sha256 and refuses on mismatch.A mailer daemon is not a user. The lesson came from the method, and the code change came from applying it to the verifier.
One address sent to twice from a log that did not know about the mailboxNothing in the product. It is recorded here as an instance of the failure the audit chain and state.json were built to prevent, arriving on the one workstream that had neither.Same reason. Nobody told us anything; a spreadsheet did.

Hypotheses under test

Stated before the interviews so they can be falsified rather than confirmed after the fact. All four are open, and open here means untested, not undecided. No evidence for or against any of them exists in this repository, because the only evidence that would count is a person's answer.

#HypothesisWhat would falsify itWhat it would take to find outWhat falls if it falls
H1The analyst, not counsel, is the person who does the change-to-action work.Two or more interviews say the real reading is done by outside counsel and the analyst only files the result.Two conversations, one utility-side and one at a firm. The corpus cannot answer it: it records who files, and this person mostly reads. Reaching them means a warm introduction or a commission-staff contact, not another cold send to a service list.ADR-001 and the whole one-user framing. The PRD gets rewritten, not patched.
H2Nobody reliably diffs versions today; it is done by reading, memory and highlighting.They show a redlining tool they already trust and use routinely for this task.One conversation, and the question has to be "show me the last one you did", not "do you use a tool". Cheapest of the four to settle and the most likely to go against us — the synthetic rehearsal already pushed back with "Compare Documents in Word, and filers do send redlines".The deterministic diff (ADR-004) stops being a wedge and becomes table stakes. The argument moves to ranking a diff rather than producing one.
H3Trust requires seeing the source passage, not a confidence score.They say a score would be enough, or that they would not check either way.Put the prototype in front of one person, ask for one task, say nothing, and watch whether they click the citation chip. Behaviour, not opinion — everybody says they would check.The depth area. Citation verification (ADR-003) stops being the trust mechanism and survives as a correctness mechanism, and the honest move is to say exactly that rather than keep selling it as trust.
H4The expensive part is interpretation and routing, not detection.They say detection is the bottleneck and interpretation is quick once they know where to look.One conversation about the last thing they missed, with a clock attached: how long between the filing landing and somebody knowing it mattered, and how much of that was finding versus deciding.The wedge moves upstream toward faster, more reliable change-finding. Most of what is built stays; what it leads with does not.

H3 is the one to watch. If trust does not require the source span, the depth area is engineering for its own sake. The response, prepared in advance: say so in the PRD, keep the verifier as a correctness mechanism because a claim that cannot show its source is still a claim that may be wrong, and stop presenting it as the reason anybody would trust the product.

The method, kept because it has never been used

Everything below is a plan. It was written before the outreach and no part of it has been run against a person. It is kept so the first real conversation does not start from nothing.

This is the old plan, not the one to run. It asks for twenty minutes and seven questions. The ten-minute call kit at the foot of this page replaces it and is what the next attempt uses. Where the two disagree, the kit wins.

Who to ask

A regulatory affairs analyst or manager at a utility, an independent power producer, or an energy retailer. Adjacent and still useful: compliance analysts in banking or pharma who track rule changes; a regulatory lawyer, discounting their answers, because their incentives differ from an in-house analyst's.

What the attempt taught about where they are. Commission service lists name every participant in a proceeding with contact details, filed as public record, and that beats any profile database — filing in a docket proves somebody does the work rather than merely holds the title. It also skews the list: of 39 targets mined that way, 13 matched the profile and 26 were commission staff, consumer advocates or outside counsel. The one untried channel with the best odds is a public-servant reader at a consumer counsel office: not the buyer and never will be, but reading the same dockets under the same clock, and far more likely to answer within a day.

The request, short enough to answer from a phone

I'm building a prototype for turning regulatory change into tracked work — dockets and orders into "here is what changed, here is what it affects, here is who reviews it." I'd value 20 minutes to hear how you do this today and where it hurts. I'm not selling anything; I want to know if I'm solving a real problem or an imagined one.

Sent 45 times. Answered none. Whether the words are the problem is not something a zero-reply run can say.

What to ask

Ask about the last time, not about preferences in general. "Walk me through the last docket you had to act on" produces facts; "would you use a tool that…" produces politeness.

AskWhat you are listening for
Walk me through the last regulatory change you had to act on. Start from how you heard about it.The trigger. An alert, a colleague, a law firm memo, a docket subscription. This tells you what the product replaces.
How did you work out what actually changed between versions?H2, directly. Whether anyone truly diffs, or reads the new version cold. If they redline by hand, that is the wedge; if they open a tool, it is not.
How do you decide whether it is material to you?The judgement the product must support rather than replace.
Who else had to look at it before anything happened?The real routing graph, against the org chart. Also the only way to test whether the obligation owner is a real second person or a design convenience.
What did you write down, and where does it live now?The system of record this has to sit alongside. If it is a spreadsheet, say so in the PRD.
When were you last burned — something missed, or acted on too early?H4, and the pain with a cost attached. This becomes the PRD's opening.
What would you need to see before you trusted a machine reading of a docket?H3, asked directly — but weigh the demo below far above the answer.

Do not describe the solution until the end. Once you do, they will be agreeable, and agreeable answers are worthless.

Then show the prototype and stay quiet

Ask for one task with no narration. Record three things: what they misread, what they ignored, and what they asked for that is not there. Silence while they struggle is the data. For H3 specifically, the measurement is whether they click the citation chip unprompted — not whether they say they would.

What this page would have to look like to stop being the weak one

Named plainly, because the gap is easier to defend when the fix is specific.

The ten-minute call kit

This section is the instrument. It is not a finding, and no part of it is evidence. It holds the script, six questions, the rules for listening, and an empty table to write a call into. Everything above this heading is an account of what has happened. Everything below it is what to do the next time somebody says yes.

A reader looking for what users said should stop at the interview table above. It is empty.

Why ten minutes and not twenty

Ten is the better ask. The old twenty went out 45 times and came back nothing — one measurement of one channel, not proof that the number was the problem. Ten fits in the gap between two meetings, so nobody has to move anything to make room, and a person who has never heard of you is agreeing to a phone call rather than to a meeting. Ten also forces the cut that twenty lets you dodge: six questions, no preamble, no demo, no slides. If the call runs long because they are enjoying it, that is theirs to offer and not yours to ask for.

Before you dial, three things and no more. Read one filing they are personally named on, so question one can name a docket instead of asking in the abstract. Have the six questions on one sheet of paper, in order, so you are not scrolling. Decide in advance which two you will drop if the clock beats you, because deciding that live is how a call ends with four polite answers and no artefact.

The script, by the clock

ClockWhat you doWhy this and not something else
0:00–0:45The open, below. Read it out. Do not improvise it.It says what you are and are not asking for before they have to guess. A stranger's first worry is that this is a sales call; answer it in the first sentence and the rest of the call is cheaper.
0:45–3:00Question 1. Then be quiet. Follow their order, not yours.The one question that buys a story rather than an answer. Everything after it is a specific you can point at inside their own account.
3:00–4:00Question 2.A count, asked early while they are still in the concrete. Late in a call this turns into an estimate.
4:00–5:15Question 3. Ask them to look, not to remember.The only question that produces an object. It is worth the awkwardness of asking somebody to open a window.
5:15–6:30Question 4.Three names, or one name three times. Either result settles something.
6:30–7:45Question 5.The trust question, asked backwards. It has to come after they trust you enough to admit they skipped something.
7:45–8:30Question 6.Short, and last, because it is about a word and words are cheap to answer once the clock is visible.
8:30–9:30Now, and only now, say what you are building. One sentence. Then two asks: would they look at it for ten minutes on another day, and who else should you be asking.No product in the first half. Once the tool is described every answer turns agreeable, so the description has to come after the answers you need. The referral is worth more than anything they said, because the channel that failed was cold.
9:30–10:00Thank them and stop. On time, even mid-sentence if you must.Stopping on time is what buys the second call. Running over is a way of telling somebody their ten minutes was a way in.

The open, word for word

Thanks for the ten minutes. Three things first, so you know what this is. I am not selling anything and there is nothing for you to look at. I am not asking for any document, any name, or anything your company has not already filed in public — if a question goes near something you cannot discuss, say pass and I will move straight on. And I am going to ask about things that already happened rather than what you would like, because what you did is the only thing I can learn from. May I take notes? And if I want to quote you in a private working document, may I use your name, or would you rather I did not?

Note what the open does not contain. No description of the product, no problem statement, no claim about the industry. The first half of this call has no product in it at all. A person who has just heard your pitch cannot then tell you what they do, because they will tell you the version of what they do that fits your pitch.

The six questions, in order

Each one names the decision it would settle. A question that cannot change anything in the build is not worth a stranger's minute, and any question added to this list later has to earn its place the same way. Two are marked can hurt: they can go against the product, and if they do, the thing that changes is large. A script that can only confirm the design is a rehearsal, not research.

#Ask, in these wordsWhat it settles
1Walk me through the last thing that landed on you that you had to do something about. Start from how you found out it existed, not from what it said.
If they generalise: which one was it, and when? Keep them on the single case.
Where the product's front door belongs. Verbatim starts when a version is handed to it. If they found out from a colleague, a law firm memo or a phone call, then the first screen is in the wrong place and ingestion is not the front door. Also the only source for the day the PRD has modelled and never observed.
2Of the last ten things that landed on you, how many had an earlier version you had already read?
If they cannot count: think of last month instead of last year.
can hurt ADR-002, which makes a change the unit of work. That holds only if most arrivals have something to compare against. If the honest answer is two, the diff wedge is a minority case, cold start stops being a gap in the product and becomes the product, and the PRD is aimed at the wrong moment.
3Open your sent mail. What was the last version comparison you sent anybody, and what was attached to it?
If nothing: when did somebody last send one to you, and what was it made with?
can hurt H2, with an object instead of an opinion. If a redline falls out of Word, or arrives from the filer already made, then the deterministic diff (ADR-004) is table stakes rather than a wedge, and the argument moves to ranking a redline rather than producing one.
4Take the last thing your company filed off the back of one of these — a comment, a correction, a compliance filing. Who wrote it, who signed it, and who decided it had to be done?
Roles are fine if names are not. Push for three, not one.
H1 and the routing model at once. If the reading is done by outside counsel and this person files the result, ADR-001 is wrong and the PRD gets rewritten rather than patched. If the decision was a meeting rather than a person, then routing to one named obligation owner is a design convenience, not the workflow.
5The last time you took somebody else's summary of a document instead of opening the document yourself — what made that all right?
If they say they never do: what was the last document you did not open at all?
H3, asked backwards. Asking what would earn trust gets a wish list; asking when trust was actually given gets a condition, and the citation chip either meets that condition or does not. Weigh what they do in a later hands-on session far above what they say here.
6When did you last write, or read, the word "material" about somebody else's change — and where did that end up?
Follow with: is that a word you would put in an email?
Whether the interpretation layer may print that word at all. The product's own screens print the word: 22 occurrences of "materiality" across app/web/templates/, 14 in change.html and 8 in project_detail.html, counted on 2026-08-05 and worth recounting rather than quoting. If it is a word they write only with counsel beside them, a screen that prints it unprompted is a liability wearing the clothes of a feature, and the label has to change.

If the clock beats you, cut in this order. Drop 6, then 5. Never drop 2 or 3: they are the two that can go against the product, and a call that keeps only the questions the product survives has told you nothing you did not already believe. Question 1 is not optional either, because it is the only one that produces their account rather than yours.

Held in reserve, if they keep talking. How long between knowing a filing had changed and knowing how it changed. What they typed out by hand because nothing gave it to them as a list. The last time they acted on something later reopened on rehearing. What happens to the document after they are finished with it — the rehearsal's own parting shot, and the one the product currently has no answer to, because Verbatim ends at a routed recommendation.

What to listen for

Four moments. Each is worth more than an answer to a question, because nobody offers them on purpose.

The workaround

The moment they describe a step they built for themselves: a spreadsheet with a tab per docket, a folder naming rule, a calendar reminder set six weeks out, a mail rule, a printout and a highlighter. A workaround is a feature somebody has already paid for with their own time, which makes it the strongest evidence available on a ten-minute call. Write the verb they use — "I pull it", "I flag it", "I run it past Dave" — and write what the thing is called. The test: did they name something with no owner, no budget and no vendor? That is a workaround, and it is a specification.

The word we do not use

Verbatim's screens print: claim, change, citation, materiality, obligation, escalation, proceeding, version, owner, project, workspace. Keep a running list of theirs. If they say erratum, service list, data request, compliance filing, redline, carve-out, tariff sheet, "run it past", "flag it", write it exactly as said and do not translate it in your notes. Translating on the way in is how a product ends up with a vocabulary nobody outside the building uses, and the note that says "they mentioned escalations" when they said "I send it to legal" has destroyed the finding before anyone could act on it. The rule: where their word and ours mean the same thing, ours is the one that is wrong.

The thing they ignore

Silence about something this product treats as central. If ten minutes about their last filing passes without them once mentioning comparing two versions, that silence outweighs any yes they give later when asked about diffing directly — the direct question supplies the idea, and people agree with ideas that are handed to them. So write down what never came up, as a list: versions, routing, deadlines, materiality, citations, draft against final. An item that never came up in ten minutes about their actual work is a claim in the PRD with nothing under it.

The restart

When they begin a sentence, stop, and start again, the second version is usually the accurate one and the first is usually the interesting one. Write both. The same goes for a hesitation before a number. And do not rescue a silence: count to five before filling it. Do not argue, do not explain how the product would handle it, do not correct them about their own job. Every one of those turns a source into an audience.

One rule over all four. Record the artefact, never the adjective. "It was a mess" is not data. "Two PDFs open side by side and a legal pad" is. If a note in the table below contains no noun you could photograph, it is an impression and should be marked as one.

The write-up, ready to fill

Three tables, empty. They exist so that writing up a call at half past ten at night is filling in cells rather than designing a document. Fill them the same day; a call written up two days later is a memory of a call. Quotes go in quotation marks and verbatim, everything else in your own words.

Table A — the call itself.

FieldCall 1Call 2Call 3
Name, and their job title in their own words
Employer: kind, and how many states
Date and start time
Minutes actually taken
Medium
How they were reached: cold, or introduced by whom
May we quote by name, anonymously, or not at all
Prototype shown: yes or no
Written up on, and by whom

Table B — the six answers. One row per question, one table per call. Copy it if a second call happens.

#The questionWhat they saidThe artefact they namedHypothesis moved, and which wayWhat changes in the build, or why nothing does
1The last thing that landed, from how they heard about it
2Of the last ten, how many had an earlier version
3Last version comparison in their sent mail
4Last thing filed: who wrote, signed, decided
5Last summary taken without opening the document
6Last time they wrote or read "material"

Table C — the four moments. Leave a row blank rather than filling it with something that nearly happened.

MomentHeard it?Their exact wordsWhat it implies for the build
A workaround they built themselves
A word they use that the product does not
Something the product treats as central that never came up
A restart or a hesitation worth keeping
If shown the prototype: what they misread
If shown the prototype: what they ignored
If shown the prototype: what they asked for that is not there

Where a filled table goes next. These three are working notes. The two tables a reviewer reads are higher up this page: the Interviews table gets a row with a real name and a real date, and What changed because of user feedback gets the last column of Table B — including the rows where the honest entry is "we heard it and changed nothing, because". A call that changes nothing is still a call, and reporting it as one is worth more than a call reported as a vindication.

What the synthetic persona can and cannot stand in for

It stands in for the interviewer's preparation and for nothing else. Rehearsing against docs/synthetic-user.html cut four questions that would have spent a real person's time producing definitions and org charts, and it produced questions 2 and 3 above — the two that can hurt most. That is the pattern worth naming: a question a synthetic cannot answer is exactly the question worth a real person's ten minutes. It can also throw up a candidate pain to go and check against a filing, which is cheap and sometimes right.

What it cannot do is answer. It was built by the same hand that built Verbatim, from a corpus that same hand chose, so its agreement is worth nothing and only its disagreement is even mildly interesting. It has no diary, so it cannot say what last Tuesday cost. It has no budget, so it cannot say what anything is worth or who signs for it. It never sees the work that produced no document, which may be most of the work. And it cannot be surprised, which is the entire reason for talking to somebody.

Nothing in this kit is evidence. Not the script, not the six questions, not the listening rules, not the empty tables. An instrument measures nothing until it is used. Evidence on this page starts when a real name and a real date sit in the interview table above; until then this product's account of its user is a model, labelled as one everywhere it appears.