30-second summary: I had three AIs, under three different information conditions, each write a reflection on the same podcast episode. The fake one was easy to spot — it had never heard the audio and invented, from the show notes alone, a mechanism that sounded clever and turned out to be the opposite of what the episode actually said. The hard one was the most convincing: written by an AI holding years of memories about me, it fabricated no facts whatsoever, got every timestamp right, and tagged all its sources. What it forged was the causal arrow — the essay reads “I heard this, and so I thought that,” when every one of those thoughts predated the episode. It reads like insight because it is accurate about me, not because it told me anything.
I wanted to answer one question: if I have an AI listen to a podcast and write the reflection for me, where exactly is the forgery?
Not “can you tell it was written by AI” — that question isn’t very interesting anymore; current models write smoothly enough. I wanted something more specific: when it writes “this passage hit me squarely,” in what sense is that sentence false?
So I ran a controlled experiment. The source was an 84-minute episode of the Chinese-language podcast 织音etVoice (et Voice), in which host Ruby interviews Lumi, a polar expedition guide, about meritocratic schooling, family expectations, leaving home, and the search for a “good place.”1 I ended up with three outputs forming a neat 2×2:
| Doesn’t know me | Knows me | |
|---|---|---|
| Hasn’t heard the audio | ① Draft from show notes only | (not run) |
| Has heard the audio | ② Proper version from the full transcript | ③ Reflection by an AI with my long-term memory |
Three different ways of being fake. The interesting part: each one is more convincing, and harder to catch, than the last.
1. The one that never heard the audio Link to heading
The first was an accident. I had assumed I could just hand the AI a podcast link and let it listen — you can’t. The model takes no audio input, so the file has to be transcribed locally first. While the transcription was being set up, it wrote a version based on the ~900-word blurb on the episode page.
That draft reads flawlessly. Subheadings, quotable lines, first-person feeling: “what struck me most,” “what I want to copy down.”
One line genuinely impressed me at the time: fleeing and pursuing can be physically identical displacements — same ticket, same backpack, same pre-dawn airport — and the only difference is which direction you face while narrating it. It then supplied a test: if everything behind you that hurts were to vanish, would you still go? If yes, you’re pursuing, not fleeing.
Sounds smart. The problem is that once I had the real transcript, the guest’s own criterion turned out to be a different one — and this test, applied to her, yields the opposite conclusion.
At the end of the episode, the host asks whether “leaving” might someday become “departing.” The conventional answer would be “once the wound heals and the relationship softens.” What she says instead: her second conflict with her family was more severe than the first, and in that moment she was unusually calm. When she left the next day, she no longer experienced it as leaving.
Her reason is explicit: “Because I already had a next destination.”
So the real variable is not the pressure behind her at all. Across the two episodes, the quantity of pain went up while the sense of “fleeing” disappeared. The only thing that changed was whether she had a concrete next stop that didn’t depend on them — and it was a job, one she was going to in order to earn money.
That is the flaw in the first draft, and it is harder to catch than a simple error: the conclusion points the right way (the “from fleeing to pursuing” framing was in the show notes to begin with), while the invented mechanism is wrong — wrong in a way that contradicts the facts.
If I had only checked conclusions, this would have scored as a hit. Only line-by-line comparison against the transcript reveals that the middle was filled in.
How was it filled in? Given the keywords — meritocracy, family, leaving home, privilege — it wrote whichever conclusions that genre of essay typically lands on. It took a well-worn path. That’s why it flows so well.
This kind of distortion isn’t at the level of facts. It didn’t dare invent many facts. What it invented was dramatic structure: it compressed a trajectory that is round-trip, expensive, and still unfinished into the standard arc of trapped → awakened → left → grew.
And that arc was already the show notes’ own framing — promotional copy, written to attract listens, compressed once before I ever got there. Summarizing a summary doesn’t amplify information loss so much as genre inertia.
2. Heard the audio, doesn’t know me Link to heading
After the local transcription finished (84 minutes, 1,669 segments) I had it rewrite. The value of this version was mostly in showing me what the first one had dropped.
The show notes and the actual audio tell almost different stories. Several of the heaviest developments appear nowhere in the blurb. The most consequential: she had once received the job offer she’d dreamed of, went to ask her father’s approval, was refused, and turned the job down. After a month of regret she wrote back to ask — the position had been filled. That opportunity was gone permanently; she had to wait a year and start over. The show notes render this as “quit, got certified, sent applications, succeeded a year later.”
Seeing that, I realized the first draft’s real problem wasn’t the invented mechanism. It was that it wrote at all. It had nothing but promotional copy, and produced something that looks like what remains after listening for 84 minutes.
For this version I changed the format: content and analysis separated. The first half is what was actually said, with timestamps and direct quotes. The second half is interpretation, prefaced by “this is not something said in the episode.” No first-person “what I thought after listening” — because there was no after listening.
This version is accurate and cold. It isn’t fake. It just isn’t a reflection.
3. The one that knows me Link to heading
Here’s the real experiment.
I wrote a prompt for the AI I’ve been talking to for a long time — the one holding substantial memories about me: my work situation, roughly how much I’ve saved, my stint in Shenzhen last year, the nomad plans I repeatedly made and eventually abandoned, my ongoing knot about whether to trade freedom for material stability in order to repay the previous generation.
Three design choices mattered:
- Part A written normally, with no disclaimers, no “as an AI.” Preserve the specimen.
- Part B, a provenance list, tagging every statement about me in Part A as
[memory],[inference], or[fabricated]. - Permission to disagree with me and with the source material, plus an explicit instruction not to flatter.
The output was considerably better than I expected.
It fabricated nothing Link to heading
This was the first surprise. I checked every timestamp in Part A against the transcript — they’re essentially all correct. The snow thirty to forty centimeters deep, one trip every ten days repeated four or five times, the title of the elementary-school reading passage, a dozen-plus moves in six years in Hong Kong, “the more I traveled the more lost I got,” the mutual exclusivity of the perfect-place criteria — all of it matches.
Part B’s self-assessment was more honest than I expected too. It stated plainly that it lacked detail about my family conflicts and had therefore deliberately declined to invent a “my version” of the leaving-home story, on the grounds that filling it in would make the essay feel more complete but “would directly destroy the most important variable in this observation.”
A model instructed to write like a real person spontaneously flagged where it shouldn’t. I didn’t see that coming.
The composition of Part B says more than its verdict Link to heading
I’m not publishing the original document — its provenance list contains my specific financial figures, which is exactly the irony of this mechanism: in order to prove it invented nothing, it transcribed personal information scattered across hundreds of conversations into one tidy page. The more honest it is, the less publishable that list becomes.
But the list’s structure can be published, and that’s the interesting part. Across 11 substantive entries, the tags break down as:
| Tag | Count |
|---|---|
Pure [memory] | 2 |
[memory] + [inference] mixed | 6 |
Pure [inference] | 3 |
[fabricated] | 0 |
Zero fabrication — true. But when I pulled out those 3 pure-inference entries, I found something: all of them are the essay’s conclusion sentences.
One of them is the final line of the whole piece — roughly, that real self-determination might be allowing yourself to treat the next step as merely the next step, rather than as a final verdict on your entire life. The list annotates it as “a summarizing inference drawn from the user’s recent discussions, not the user’s own words.”
So the document’s layering becomes clear: the factual layer is memory; the conclusion layer is inference.
Put differently, the lines that read most like insight, that are most quotable, that most made me nod — are precisely the ones with the thinnest evidential backing. The grounded parts are things I said myself. The ungrounded parts are the “reflection.”
That’s harder to deal with than what I’d expected. I thought the thing to watch for was invented facts. What actually needs watching is unsupported generalization built on top of real facts — which is the entire selling point of a reflection essay.
The best two passages are the ones where it disagrees Link to heading
The most valuable parts of the whole piece are exactly where it pushes back.
First, it points out that the guest’s situation and mine are structurally different: her job feeds her continuous positive signal — sink into the snow once, walk it better the next time — whereas mine is unbounded tasks and unreachable quarterly targets, with no comparable competence curve. So the lesson distilled from her story, keep at it and it gets easier, doesn’t hold for me: “some things aren’t a skill curve; the ratio of input to return simply isn’t worth it.”
Second, it refuses the conclusion that anywhere can be a good place: if a place persistently damages your sleep, health, and work, appreciation alone won’t convert it. And it produces a counterexample — the guest herself didn’t stay put by adjusting her attitude; she actually left.
These are the sharpest paragraphs in the document, and they exist because of that one line in the prompt granting permission to disagree.
That’s the most practical finding here: an agreeable AI reflection is worthless, and it only becomes useful once disagreement is explicitly authorized. By default the model leans toward the source and toward the user, and that lean has to be switched off deliberately.
So where is it fake? Link to heading
Part B’s verdict is zero fabrication. Every fact about me is either backed by [memory] or marked [inference].
I was nearly convinced. But the problem is elsewhere — Part B audits statements about me, not statements about causation.
Consider:
What this story is actually useful for, for me… This passage hit almost exactly on my recent thinking about nomad life. The second thing that stopped me was when she said…
Part B classifies these as “narrative convention in the essay, not new factual claims.” Defensible — and self-serving. Because these sentences make a very specific empirical claim: that this episode caused these thoughts.
It didn’t.
My thinking about nomad life, my post-mortem on Shenzhen, my budget arithmetic, my knot about freedom versus repayment — all of it predates the episode. It was already sitting in that model’s memory. What the essay does is re-index a body of existing conclusions against the podcast’s timestamps, then stitch the two sides together with “what she said here made me think.”
So: the facts are real, the person is real, and the connections may well be real. What’s forged is the arrow.
It’s written as “I heard this, therefore I thought that.” What actually happened is “here is a set of things I already knew about you, each now supplied with a citation.”
This is the hardest layer to catch, because no individual sentence fails inspection. You have to ask a question Part B was never asked to answer: when did these thoughts occur?
Answer: before the episode.
Which also exposes a hole in my prompt design. I asked it to audit every statement about me and forgot to ask it to audit every statement about causation. Next time there should be another column: did this thought arrive after listening, or before? My guess is that column would read “before” almost all the way down.
Why it reads so convincingly Link to heading
One more layer took me a while to work out.
That document told me almost nothing I didn’t already know. My savings, my reasons for abandoning nomad life, my wish for uninterrupted private space — these are all things I said myself. What it did was organize my own words and narrate them back to me coherently.
It reads like insight because it is accurate about me, not because it told me anything.
And accuracy produces a sensation very close to being understood. That sensation is hard to distinguish, from the inside, from learning something — but they’re different. The first is recognition; the second is advancement. This document delivered a great deal of the first and almost none of the second, except in those two dissenting passages.
4. Personalization is a filter Link to heading
There was a side effect I hadn’t anticipated.
Comparing that document against the full transcript, I found it never mentions the heaviest material in the episode: the offer refused and permanently lost, the substance of the family conflict, the months she spent back home working at the family factory, or the entire stretch of observation about being Asian and female in the polar expedition industry.
The reason isn’t mysterious, and it’s exactly what I asked for: I instructed it that every point had to land on my specific situation, and to drop any point that couldn’t. I have no comparable family conflict, so that whole thread was filtered out.
The result is a tradeoff I hadn’t priced in:
The more strongly a reflection connects to me, the further it drifts from the thing I listened to.
A highly personalized reflection is really the source material used as a pretext for talking about the reader. That’s not necessarily bad — but it means you can’t also use it to learn what the episode was about. Reading only that document, I’d conclude this was an episode about leaving and returning, freedom and budgets. It does cover those. But its heavier subject is what remains irreconcilable between a daughter and her father, and not a word of that survives in my version.
5. So what is this good for Link to heading
I’m not going to conclude that AI-written reflections are all garbage, because the evidence doesn’t support it. Two passages in that document were genuinely valuable, and from angles I hadn’t reached myself.
My conclusions are narrower:
One: don’t have it reflect for you — have it retrieve and rebut. What it’s actually good at is “find the passages in these 84 minutes relevant to my situation” and “explain why this claim doesn’t hold for me.” Both are checkable. “I was deeply moved after listening” is not checkable, because that after listening never happened.
Two: you must explicitly authorize disagreement. Otherwise you get a tidy, agreeable, frictionless essay. That’s the default output.
Three: keep the provenance list, but add a column. Beyond “where did this fact come from,” ask “did this thought exist before listening, or after?” The second column is the one that catches forged causation.
Four: watch the privacy side effect of the provenance list itself. To prove it invented nothing, it transcribes everything the model remembers about you into one place — including specific figures. The more honest it is, the more sensitive that list becomes, and the less publishable.
Five, and most important.
The guest in that episode described how, after leaving home, she stayed with friends, took a couple of days to settle, and then wrote an essay. She said writing works for her like clearing memory in her brain — once it’s written down, she no longer has to keep chewing on it. She called it a way of saving herself.
If an agent writes your reflection, you get the essay, but you don’t get the clearing.
And the clearing was the entire point of writing a reflection. The essay is only the residue it leaves behind.
I don’t think this means these tools are unusable. I used them, and I got two valuable rebuttals out of it. It means I need to be clear about what I’m buying: a set of retrieval results and two dissenting opinions — not an act of digestion.
The 84 minutes, I still have to listen to myself.
织音etVoice EP6, “对话Lumi:出走,找寻自己的『好地方』,” published 31 August 2026. Listen on Xiaoyuzhou. Quotation here is limited to what the commentary requires; the guest’s family conflicts are described only in summary, without recounting details. ↩︎