Hindsight: Vectorize's Open-Source AI Agent Memory (2026)
Hindsight is an open-source (MIT) agent memory system by Vectorize built to create agents that learn over time, not just recall chat history — it hits state-of-the-art scores on the LongMemEval benchmark. Updated Sep 2026.
Hindsight is an agent memory system. While most similar tools focus on recalling conversation history, Hindsight is built around making agents that learn over time. It's developed by Vectorize (vectorize-io), licensed under MIT, implemented in Python, and is described in the official README as eliminating the shortcomings of alternatives such as RAG and knowledge graphs while delivering state-of-the-art performance on long-term memory tasks.
As of this writing (September 28, 2026), the repository has roughly 39,777 stars and 5,295 forks. It ranked #3 on GitHub's weekly trending list with +11,089 stars in a single week (#2 was Paperclip, an agent-team-management tool). The latest release is v0.10.1, shipped 2026-09-21; the repository was created on October 30, 2025. This article covers Hindsight's core concepts, installation, basic usage, and how it compares to other tools, based only on the official README and GitHub repository.
What Is Hindsight? Built to Learn, Not Just Remember
The README states: "Most agent memory systems focus on recalling conversation history. Hindsight is focused on making agents that learn, not just remember." It reports state-of-the-art results on LongMemEval, a widely used benchmark for long-term memory tasks, and says this has been independently reproduced by research collaborators at Virginia Tech's Sanghani Center for Artificial Intelligence and Data Analytics and The Washington Post. It's reportedly used in production at Fortune 500 enterprises and a growing number of AI startups.

Core Concepts — Retain, Recall, Reflect, and Four Memory Types
Hindsight says it organizes memory in a way closer to human memory, storing it in an isolated store called a bank. New memories are pushed into either the World Facts pathway (facts about the world, e.g. "the stove gets hot") or the Experiences pathway (the agent's own experiences, e.g. "I touched the stove and it really hurt"), and represented as a combination of entities, relationships, and time series with sparse/dense vector representations. From there, Observations — deduplicated, evidence-backed beliefs consolidated from many memories — are formed, and Mental Models synthesize a learned understanding from Observations and facts.
| Operation | What it does |
|---|---|
| Retain | Pushes new information into Hindsight. An LLM extracts facts, temporal data, entities, and relationships, then normalizes them into canonical entities, time series, and search indexes |
| Recall | Retrieves memories. Runs four retrieval strategies in parallel — semantic (vector), keyword (BM25), graph (entity/temporal/causal links), and temporal range filtering — merges results with reciprocal rank fusion and cross-encoder reranking, and trims to fit a token budget |
| Reflect | Performs deeper analysis over existing memories, forming new connections or answering questions that need reasoning rather than lookup — e.g. an AI project manager assessing risk, or a sales agent reflecting on why some outreach gets responses |
Retained facts aren't just piled up — related facts are consolidated in the background into Observations, each carrying its supporting evidence (exact quotes, a proof count) and refined rather than overwritten as new evidence arrives, so new information strengthens, weakens, or extends an existing belief. A step further is the Mental Model: a standing answer to a question about a bank ("what are this user's preferences?"). Define the question once, and Hindsight writes and keeps rewriting the answer in the background — reading one is a plain database read, no LLM call needed. Knowledge Pages hide the mechanics of a mental model, presenting it as a folder-organized, searchable Markdown document instead.
A bank is an isolated memory store — typically one per user, agent, or project — with strict isolation and no cross-bank leakage. Banks carry "disposition traits" (skepticism, literalism, empathy) that shape how reflect reasons over their memories. Input language is detected and preserved by default (multilingual), so entities keep their native script instead of being translated, and an opt-in Memory Defense feature scans every retain call against 45 patterns for secrets and PII, redacting or blocking matches.
Installation and Pricing
| Item | Detail |
|---|---|
| Developer | Vectorize (vectorize-io) |
| License / price | MIT, open source, free to self-host. Managed Hindsight Cloud is usage-based with free starter credits, no fixed monthly or per-seat fee |
| Implementation | Python server (hindsight-api); clients available for Python, Node.js, Go, and CLI |
| Supported LLMs | 25+ hosted providers (OpenAI, Anthropic, Gemini, Groq, Bedrock, Vertex AI, MiniMax, DeepSeek, and more), fully local Ollama/LM Studio/llama.cpp, any OpenAI-compatible endpoint, and gateways like LiteLLM. Existing ChatGPT Plus/Pro, Claude Pro/Max, Cursor, and GitHub Copilot subscriptions work without an API key |
| Supported platforms | Linux (x86_64, ARM64), macOS (Apple Silicon, Intel), Windows (x86_64) — via Docker, pip, or an embedded DB (pg0); Intel Macs need hindsight-all-slim for the pip install |
| Storage | PostgreSQL + pgvector, or Oracle AI Database 23ai with full feature parity |
| Latest version | v0.10.1 (2026-09-21) |
| References | LongMemEval benchmark paper (arXiv:2512.12818), a public benchmarks page |
The server can be started via Docker, pip (bare metal), Helm (Kubernetes), or an embedded DB with no server at all. Docker is the quickest path.
export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 \
-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
-v hindsight-data:/home/hindsight/.pg0 \
ghcr.io/vectorize-io/hindsight:latest
# API: http://localhost:8888
# UI: http://localhost:9999For a bare-metal pip install:
pip install hindsight-api
export HINDSIGHT_API_LLM_API_KEY=sk-xxx
hindsight-apiHow to Use It — The Shortest Path
Once the server is running, install a client and connect.
pip install hindsight-client -U
from hindsight_client import Hindsight
client = Hindsight(base_url="http://localhost:8888")
# Retain: store information
client.retain(bank_id="my-bank", content="Alice works at Google as a software engineer")
# Recall: search memories
client.recall(bank_id="my-bank", query="What does Alice do?")
# Reflect: generate a disposition-aware response
client.reflect(bank_id="my-bank", query="Tell me about Alice")For the fastest way to add memory to an existing agent, use the LLM Wrapper, which swaps in for your LLM client. It defaults to Hindsight Cloud; pass hindsight_api_url to point it at a self-hosted server instead.
pip install hindsight-litellm
from openai import OpenAI
from hindsight_litellm import wrap_openai
client = wrap_openai(
OpenAI(),
bank_id="user-123",
hindsight_api_url="http://localhost:8888",
)
# Hindsight recalls relevant memories before the call
# and retains the conversation after it
response = client.chat.completions.create(
model="gpt-5-mini",
messages=[{"role": "user", "content": "What do you know about me?"}],
)wrap_anthropic() does the same for the Anthropic SDK, and since LiteLLM sits underneath, the same integration covers 100+ models. For explicit control over when memories are stored and recalled, use the SDKs or REST API directly. There are 60+ integrations, and coding agents (Claude Code, Codex, Cursor, GitHub Copilot) and agent frameworks (LangGraph, CrewAI, OpenAI Agents SDK) mostly need no code changes. A dedicated package builds automatic project memory for coding agents from git history and past sessions.
npx @vectorize-io/hindsight-coding-agents install all # every detected agent, wired natively
npx @vectorize-io/hindsight-coding-agents install claude-code # or just oneEvery server also ships a built-in MCP endpoint per bank at http://localhost:8888/mcp/{bank_id}/, so any MCP client can call retain, recall, and reflect as tools. For MCP background, see the MCP roadmap breakdown.
How It Differs From Existing Tools
| Hindsight | mem0 | Letta | Zep (Graphiti) | |
|---|---|---|---|---|
| Focus | A memory layer for "agents that learn": retain/recall/reflect over four memory representations | A memory layer for personalized AI, holding context across User, Session, and Agent levels | A framework for "stateful agents" (formerly MemGPT) that learn by rewriting their own memory blocks over time | A temporal knowledge graph that tracks how facts change over time. Graphiti is the open-source core framework; Zep is the commercial managed product built on it |
| Retrieval | Four strategies in parallel (semantic, keyword/BM25, graph, temporal), merged via RRF + reranking | "Multi-signal retrieval": semantic, BM25 keyword, and entity matching scored in parallel and fused (per the April 2026 algorithm update) | No explicit multi-strategy retrieval described in the README; the agent rewrites its own memory blocks, and /search searches across all messages and agents | Hybrid retrieval combining semantic, BM25 keyword, and graph traversal, plus bi-temporal tracking of when facts became true or were superseded |
| Benchmarks | SOTA on LongMemEval, independently reproduced by Virginia Tech and The Washington Post; public benchmarks page compares vendors | Self-reported: LoCoMo 92.5, LongMemEval 94.4 (April 2026 algorithm update; scores reflect the managed Platform, with mem0 noting OSS users should expect directionally similar but not identical numbers) | — (no LongMemEval/LoCoMo figures published in the README) | — (no LongMemEval/LoCoMo figures published in the README; Zep's paper reports its own DMR benchmark instead) |
| License | MIT (free self-hosted) + usage-based Hindsight Cloud | Apache 2.0. Library and self-hosted server are free; Mem0 Platform is a paid managed tier | Apache 2.0 (Letta Code). CLI/self-hosted is free; the default Letta Cloud backend requires sign-in for some features | Graphiti (the OSS core) is Apache 2.0, free, self-hosted only. Zep itself is a proprietary commercial product (fully managed or deployed in your own cloud) |
| Deployment | Docker / pip / Helm / embedded DB / managed cloud | Library (pip/npm) / self-hosted server (Docker Compose) / cloud platform (app.mem0.ai) | Install the CLI via npm and run locally or via Letta Cloud; desktop app, browser, and Slack/Telegram/Discord integrations also available | Graphiti requires bringing your own graph DB (Neo4j / FalkorDB / Amazon Neptune, etc.) and self-hosting it; Zep offers fully managed or in-your-cloud deployment |
mem0, Letta, and Zep are all well-known open-source projects in agent memory and agent state management, and the table above is built from each project's own official README and GitHub repository. mem0 is Apache 2.0, offered as a library and self-hosted server plus a paid managed Platform, and self-reports LoCoMo 92.5 / LongMemEval 94.4 on its April 2026 extraction algorithm (scores from the managed Platform). Letta descends from MemGPT and takes a "stateful agent" approach where the agent itself rewrites its own memory blocks to learn, emphasizing self-editing context management over an explicit multi-strategy retrieval API. Zep is a commercial managed product built around a temporal knowledge graph, and its open-source core, Graphiti (Apache 2.0), is a self-hosted-only framework that requires bringing your own graph database such as Neo4j. What the README emphasizes as Hindsight's distinctive trait is combining multiple representations — World Facts, Experiences, Observations, and Mental Models — rather than plain vector search or a knowledge graph alone, continuously refining evidence-backed beliefs, and publishing results on LongMemEval, a benchmark independently reproducible by third parties. The official RAG vs Hindsight page explains the RAG comparison in more depth. For a RAG platform build example, see the full WeKnora guide.
Things to Watch For, and Who It's For
- Operations that call an LLM (fact extraction during retain, synthesizing Observations and Mental Models) incur separate API costs from whichever LLM provider you configure
- The README explicitly says Hindsight "may be overkill" for simple AI workflows like those built with n8n, and is aimed at agents that handle open-ended tasks, change behavior from feedback, and learn to perform complex work — closer to an "AI employee"
- Installing via pip on an Intel Mac (x86_64) requires hindsight-all-slim rather than the standard hindsight-all
- Memory Defense is opt-in, so it must be explicitly enabled for use cases that need to protect secrets or PII
- Production self-hosting comes with a full set of operational tooling — Prometheus monitoring, an admin CLI, webhooks — but each of these needs its own setup
It's a good fit for anything you want to keep getting smarter across sessions: chatbots that remember per-user preferences, coding agents that build up project understanding over time, or sales/support agents reflecting on patterns in customer interactions. For simple Q&A workflows with no need for long-term memory, it isn't necessarily required.
FAQ
Is Hindsight free to use?
The server itself is open source under MIT and free to self-host. The managed Hindsight Cloud option is usage-based with free starter credits, and the official site states there's no fixed monthly or per-seat fee.
Which LLM providers does it support?
25+ hosted providers including OpenAI, Anthropic, Gemini, Groq, Bedrock, and Vertex AI, plus fully local options like Ollama, LM Studio, and llama.cpp, OpenAI-compatible endpoints, and gateways such as LiteLLM. Existing ChatGPT Plus/Pro, Claude Pro/Max, Cursor, and GitHub Copilot subscriptions reportedly work without a separate API key.
How is this different from RAG?
According to Vectorize, where RAG typically relies on plain vector search or a knowledge graph, Hindsight combines multiple memory representations (World Facts, Experiences, Observations, Mental Models) and continuously refines evidence-backed beliefs. See the official RAG comparison page for details.
Does it work with coding agents like Claude Code or Cursor?
Yes. Installing the dedicated @vectorize-io/hindsight-coding-agents package automatically builds project memory from a repo's git history and past sessions, along with knowledge pages covering architecture and conventions, loaded in as the agent starts working.
Does it support languages other than English, like Japanese?
The official docs state it is multilingual by default: input language is detected and preserved, and entities keep their native script rather than being translated (the docs give a Chinese-name example that stays as written).
Summary
Hindsight is a memory layer for "agents that learn," not just a store to search chat history — it keeps refining evidence-backed beliefs and grows standing answers (Mental Models) over time. Its combination of a third-party-reproduced SOTA score on LongMemEval, a simple retain/recall/reflect API, and 60+ ready-made integrations is what stands out.
That said, any operation involving an LLM call adds provider costs, and the README itself notes it can be overkill for simple workflows. Starting with a single small bank via Docker or the embedded DB, confirming how retain and recall behave, and then turning on production features like Mental Models and Memory Defense in stages looks like the reasonable path forward.
Sources
- https://github.com/vectorize-io/hindsight
- https://hindsight.vectorize.io
- https://hindsight.vectorize.io/developer/rag-vs-hindsight
- https://benchmarks.hindsight.vectorize.io
- https://arxiv.org/abs/2512.12818
- https://github.com/vectorize-io/hindsight/releases/tag/v0.10.1
- https://github.com/mem0ai/mem0
- https://github.com/letta-ai/letta-code
- https://github.com/getzep/graphiti
- https://www.getzep.com
- https://github.com/trending?since=weekly
Feel free to contact us
Contact Us