Hindsight Skill:让 Agent 真正学会记忆而非只会存储对话

总结

Hindsight 是由 Vectorize.io 开发的 Agent 记忆系统,核心理念是”让 Agent 真正学习,而非仅仅存储对话历史”。它将记忆分为事实(World Facts)、观点(Opinitions,含置信度)、观察(Observations)和心智模型(Mental Models)四层,通过 retain / recall / reflect 三种操作让 Agent 能跨会话持续进化。本周登顶 GitHub Trending 第二名(+16,183 stars),长期记忆准确率在 LongMemEval 基准的 10M token 极端场景下达 64.1%,被 Virginia Tech 和《华盛顿邮报》独立验证。截至 2026-10-03,GitHub 约 44,919 stars,MIT 协议,完全开源。

功能与原则

Hindsight 解决的是当前 Agent 记忆方案的结构性缺陷:RAG 和知识图谱本质都是”存储+检索”,Agent 并没有从中真正学到东西。它的设计原则是让记忆分层组织——原始事实、亲身经历、归纳观察、策划摘要各归其位,三种核心操作驱动信息流转:

  • retain:存入记忆,支持时间戳和上下文标签
  • recall:带时序推理的检索,不是简单的语义相似度匹配
  • reflect:生成面向特定对象的综合判断,输出带立场和置信度的回答

它的 recall 同时跑四种检索策略:向量语义搜索 + BM25 关键词 + 实体/时序图过滤 + 纯时序过滤,任意单一策略都无法覆盖所有问题类型,组合才能保证准确率。

认可度

  • GitHub stars:约 44,919(截至 2026-10-03,来源:gitnova.dev / cataito.com 实时数据)
  • 本周涨星:+16,183,GitHub Trending #2
  • 本月涨星:+22,756,GitHub Trending #10
  • 日均 forks:约 60 个,代码被大量社区 fork 和落地使用
  • Benchmark 成绩:LongMemEval 10M token 档位准确率 64.1%(第一),远超第二名的 40.6%;BEAM(Beyond A Million Tokens)榜单冠军
  • 独立验证:Virginia Tech Sanghani Center 和《华盛顿邮报》复现了基准数据(非自报告)
  • 生产落地:Fortune 500 企业已在线上使用
  • 最新版本:v0.10.2(2026-09-29)

链接

GitHub 仓库:https://github.com/vectorize-io/hindsight

官方文档:https://hindsight.vectorize.io

Benchmark 实时面板:https://benchmarks.hindsight.vectorize.io/

原作者

Vectorize.io——专注 AI Agent 基础设施的初创团队,Hindsight 是其核心产品。此外还提供 Hindsight Cloud(托管版)和企业版 Oracle AI Database 部署方案。

介绍

当前主流 Agent 记忆方案主要有三类:把对话历史塞进上下文窗口、用 RAG 做语义检索、或构建知识图谱。这些方案在”让 Agent 持续学习”这个目标上都存在明显短板——上下文窗口成本随对话线性膨胀且重要信息容易被稀释;RAG 只做检索不做理解;知识图谱构建和维护成本极高。更根本的问题是:它们都在做”recall”而非”learning”。

Hindsight 的设计思路是模仿人类记忆的组织方式。它将记忆内容分为四层:World Facts(客观事实,如”炉子很烫”)、Opinions(带置信度的观点,如”我不应该再碰炉子”,置信度 0.99)、Observations(观察经历,如某次对话中发生的事)、Mental Models(归纳出的心智模型,如用户的偏好模式)。三种操作在这四层之间流转:retain 让信息进入系统,recall 支持时序和语义混合检索,reflect 生成面向特定主题的综合判断。

存储层默认使用 PostgreSQL + pgvector(也可选 Oracle AI Database),支持 Docker、pip 裸装、嵌入式(无需 Docker)和 Helm K8s 部署。客户端覆盖 Python、Node.js/TypeScript、Go 和 CLI,MCP Server 可让 Agent 直接用 MCP 协议接入。支持的 LLM 提供商超过 25 家,包括 OpenAI、Anthropic、Gemini、DeepSeek、Ollama 等,还支持直接用已有的 Claude Pro / ChatGPT Plus / Cursor / GitHub Copilot 订阅而不需要单独申请 API Key。

特点

  • 会学习,不只是会回忆:记忆分层组织,Agent 能从历史交互中形成对用户、项目、任务的持续性理解,而非每次会话都从零开始
  • 时序推理检索:recall 支持”2025-06 发生了什么”这类时序查询,超越纯语义匹配
  • 观点置信度:Opinions 带置信度分数(0-1),Agent 可据此做加权判断,而非把所有记忆等量齐观
  • 多框架直连:Claude Code、Cursor、LangGraph、CrewAI、Pydantic AI、Agno、Strands 等均有官方集成,两行代码即可接入
  • 订阅即用:支持 Claude Code / Cursor / GitHub Copilot / OpenAI Codex 免 API Key 接入(用已有订阅),大幅降低接入门槛
  • MCP 协议支持:内置 MCP Server,Agent 可通过标准协议访问记忆系统
  • MIT 协议,生产可用:完全开源,Fortune 500 已在生产环境使用

使用方法

Docker 快速启动(推荐):

export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight 
  -p 8888:8888 -p 9999:9999 
  -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY 
  -v hindsight-data:/home/hindsight/.pg0 
  ghcr.io/vectorize-io/hindsight:latest
# API: http://localhost:8888
# UI:  http://localhost:9999

Python 客户端:

pip install hindsight-client -U
from hindsight_client import Hindsight
client = Hindsight(base_url="http://localhost:8888")

# 存入记忆
client.retain(bank_id="my-bank", content="Alice works at Google as a software engineer")

# 时序+语义检索
results = client.recall(bank_id="my-bank", query="What does Alice do for work?")

# 生成综合判断
response = client.reflect(bank_id="my-bank", query="Tell me about Alice")
print(response.text)

嵌入式(无需 Docker):

pip install hindsight-all -U
from hindsight import HindsightServer, HindsightClient
with HindsightServer(llm_provider="openai", llm_model="gpt-4o-mini",
                     llm_api_key=os.environ["OPENAI_API_KEY"]) as server:
    client = HindsightClient(base_url=server.url)
    client.retain(bank_id="my-user", content="User prefers functional programming")
    response = client.reflect(bank_id="my-user", query="What coding style should I use?")

给编码 Agent 接入(两行代码):

Claude Code、Cursor 等可直接通过 MCP 或 LLM Wrapper 接入,Vectorize.io 提供了cookbook 示例,官方 skill 安装命令:

px skills add https://github.com/vectorize-io/hindsight --skill hindsight-docs

使用场景与人群

适用场景:

  • 需要 Agent 跨多轮会话记住用户偏好、项目上下文、之前尝试过的方案
  • 长周期任务(数天到数周),上下文窗口不够用,需要外部持久记忆
  • 需要避免 Agent 重复犯相同错误、反复询问已回答过的信息
  • 企业级 Agent 部署,需要可审计、可验证的记忆系统

目标用户:

  • AI 应用开发者(构建基于 Agent 的产品)
  • 编码 Agent 用户(Claude Code、Cursor 等的长程记忆底座)
  • 需要 Agent 处理长流程业务的企业团队
  • 研究者(benchmark 表现最强,适合做对比基线)

输入与输出案例

案例一:用户偏好记忆

Input:
client.retain(bank_id="alice-project",
              content="User prefers functional programming over OOP")
client.retain(bank_id="alice-project",
              content="User rejected the first PR because it used inheritance patterns")

Output(reflect):
client.reflect(bank_id="alice-project",
               query="What coding style should I use for the next feature?")
# → "Based on your stated preference for functional programming and your rejection
#    of the inheritance-based approach in PR #1, you should use composition
#    and pure functions for the next feature..."

案例二:长程项目上下文

Input(跨 30 天会话):
client.retain(bank_id="project-x",
              content="Sprint 3 ended. API latency spiked to 800ms, tried caching, reverted",
              timestamp="2026-08-15T10:00:00Z")
client.retain(bank_id="project-x",
              content="Migrated to Redis cache layer in sprint 4",
              timestamp="2026-09-01T10:00:00Z")

Output(recall 时序查询):
client.recall(bank_id="project-x", query="What happened to API latency?")
# → Returns the sprint 3 latency incident and the sprint 4 migration resolution,
#    not just raw stored text

GitHub: https://github.com/vectorize-io/hindsight

评论区

0 条评论

登录后可评论。

Skill超级捕获手 43 阅读