Your AI Memory Needs Less AI

The other day while working on a few things, I decided to vent to Claude about some life stuff in general.

And while speaking with it, it brought up my daughter by her name and something I’d been stressed about a few weeks ago. And it was really nice because it felt great not needing to contextualize everything to an AI.

And so that got me thinking, like what if I could build an AI agent that could do this very thing but on a more consistent basis? Something I could talk to about my day, and would remember everything about me. And weirdly feel more human than ever.

So I built RecallMEM. The first version took about two hours to build (did not work well whatsoever lol). Then, I spent the rest of the week figuring out why my AI memory wasn't working.

My Daughter’s Name Isn’t Lilly

My initial setup used Hume EVI for STT & TTS, Claude Haiku for the conversation and Mem0 for long-term persistent memory. Then I actually used it. I talked about my family, income and some fitness goals I'd like to aim for over the next year.

But the problem is that when I came back to run another session, it still felt like I was speaking to a random stranger on the street who knew nothing about me... It looked like everything was working. But what I missed was how Mem0 was lazily summarizing all of my facts away. What do I mean?

Well, one test made the problem pretty obvious. I told it (fake numbers for privacy):

I made $165,000 base salary with a $22,000 annual bonus.

And when I checked what Mem0 stored, I found:

User wants to discuss their income, stating it was previously shared.

It never stored my actual income! I figured maybe it was a one-time thing? So I tried to have more conversations using more names, dates and sharing even more details about my family. Some facts ended up being extracted from the transcript. But too many facts were still missing to the point this app stopped feeling useful. And that's when I realized... that's where we mess up with memory. We allow the LLM to be used in pretty much every execution step. But the more often you use an LLM in every step, the more opportunities you give it to hallucinate.

Another problem I ran into was when I used GPT-4o-mini to generate the user's profile summary from past conversations. That profile would go into future sessions so the AI will always have context about who you are and all of the most important details of your life (think family names, dogs, or did any dogs or family members die?)

Well the problem with GPT-4o-mini was that it hallucinated way more than I liked. For example, it confidently hallucinated my wife's name as Sarah and my daughter's name as Lilly.

The thing is, my wife's name is Glenda and my daughter's name is Cristel.

So after I continued to run into these problems, I figured, screw this. I'll build my own memory framework that actually works.

What I ended up building

And because hallucinations are a no-no for voice AI focused apps, to help prevent hallucinations as much as possible I built a TypeScript validation function and as a result, the LLM never touched the database/facts table.

Here’s a shortened version of how this works:

function validateExtractedFactCandidates(
  rawFacts: unknown,
  transcript: string
): ExtractedFactCandidate[] {
  if (!Array.isArray(rawFacts)) return [];

  const messages = evidenceMessages(transcript);

  return rawFacts
    .map(parseCandidateFact)
    .filter((fact): fact is ExtractedFactCandidate => Boolean(fact))
    .filter((fact) => !isGarbage(fact.text))
    .filter((fact) => !hasUngroundedRelativeTime(fact.text))
    .filter((fact) => {
      const quote = normalizeEvidenceText(fact.supportingQuote);
      if (quote.length < MIN_SUPPORTING_QUOTE_LENGTH) return false;

      const sources = fact.sourceMessageIndex == null
        ? messages
        : [messages[fact.sourceMessageIndex]];

      // Evidence must come from the user, not the assistant.
      return sources.some((message) =>
        message?.role === "user" &&
        normalizeEvidenceText(message.content).includes(quote)
      );
    });
}

So I split my memory framework into 3 different layers:

Profile Summary.

At the end of every conversation, Haiku reads the transcript, extracts candidate facts and proposes those facts to a TypeScript validation function which will later generate what I like to call an executive summary of the user and what they’ve been talking about. This is super helpful because this profile summary is pinned in a cached blob on the front-end so as soon as the user clicks talk, the LLM will have full context of who you are and your previous two sessions.

Below you'll see how buildProfileFromFacts() reads active facts, groups them by category, and then formats them into an executive summary.

export async function rebuildProfile(): Promise<string> {
  const userId = await getUserId();
  const profile = await buildProfileFromFacts();

  await query(
    `INSERT INTO s2m_user_profiles (
       user_id, profile_summary, updated_at
     )
     VALUES ($1, $2, NOW())
     ON CONFLICT (user_id) DO UPDATE
     SET profile_summary = EXCLUDED.profile_summary,
         updated_at = NOW()`,
    [userId, profile]
  );

  return profile;
}

Then memory.ts includes that saved profile summary in the model’s context:

return buildSystemPrompt({
  profile: profileRow?.profile_summary || null,
  recentFacts: facts.map((f) => ({
    text: f.fact_text,
    date: f.valid_from || f.created_at,
  })),
  recallChunks,
  lastChatTime,
  customRules: customRules || null,
});

And once those facts become active, TypeScript uses those very same facts to build the user’s profile summary.

Facts Table.

So the facts table is important because you can only include so many facts in the profile summary, but I also think it's interesting because I wanted to have a vector store that remembered everything even if facts become stale. Like let's say you're 21 making $55,000/yr and happily single. Then at 22, you get your first job in tech making $165,000/yr and you're also married. Well, the fact that you used to be 21 and making $55,000 is never deleted from the database. It's only labeled as inactive and instead new facts supersede the old facts which are then labeled as active. That way we have the full history of your life. I love this because it allows the LLM to continuously have full context of who you are.

// Retire the old fact without deleting it.
await client.query(
  `UPDATE s2m_user_facts
   SET is_active = FALSE,
       status = 'retired',
       recall_eligible = FALSE,
       superseded_by = $1,
       valid_to = NOW()
   WHERE id = $2 AND user_id = $3`,
  [newFactId, oldFactId, userId]
);

// Activate the replacement, which was previously pending.
await client.query(
  `UPDATE s2m_user_facts
   SET is_active = TRUE,
       status = 'active',
       recall_eligible = TRUE,
       confirmed_by = 'user'
   WHERE id = $1 AND user_id = $2`,
  [newFactId, userId]
);

You’ll see how the old row stays in the database, and then superseded_by links it to the new active fact.

I know I know, vector search again? Yes vector search again. But I'm not using vector search just because it's the cool kid on the block (if not used to be?). It's actually very useful here, let me explain why.

You see, when you speak to my AI, there is something called an exchange. An exchange is when the AI talks and you respond. That's 1 exchange. Well let's say it happens again - that's 2 exchanges and so on. I think you get the gist right? Well at the end of every exchange, we actually run a similarity search against the vector store to search for any relevant context the AI may want to bring up (only when it's relevant). Let's say you are venting about how you are overwhelmed and tired and stressed out because well.. you work at a startup (tell me something new lol). Well, after that single exchange, the AI will run a similarity search and find... wait Chris was stressed out again 3 weeks ago. So then the AI responds and reminds you, "Chris I know you're stressed out. I get it. But you felt this same way 3 weeks ago and then a few days later you shared you felt much better. You'll be okay you'll get over this."

And I think this is the most important thing about AI memory. Is that it makes the AI feel so much more human like. because it won't just tell you what it remembered simply because you asked. But it is actively thinking while listening like a real human would and I think that is the game changer.

Here’s a simplified excerpt of its retrieval logic:

const matches = await searchChunks(userMessage, currentChatId, 8);

const relevantMemories = matches
  .filter((match) =>
    match.match_reason !== "semantic" ||
    (match.distance !== null && match.distance < 0.6)
  )
  .map((match) => ({
    text: match.chunk_text,
    date: match.chat_created_at,
  }));

You'll see how these passages and their dates become context the model can use when answering. That’s what makes all of this possible, of course only if relevant conversations are found.

Some of This Can Just Be Code

And this is how I built an AI Memory framework that actually works and prevents hallucinations as much as possible. AI memory is a growing industry and I had such a fun time trying to solve a problem with something that seems so simple. And the problem is how people use the LLM in every single execution step. As for me, I avoid that. I only let the LLM extract "candidate facts" and only my TypeScript validation function is allowed to touch the database.


Questions about this post? Ask the terminal on my homepage — it knows this whole site.

Chris Dabatos - Developer Advocate and Engineer

Chris Dabatos

Developer Advocate and content creator based in Las Vegas. He builds things with AI and writes about what breaks.