AI-Chain

TencentDB Agent Memory: From Flat Vector Piles to Traceable Layered Memory

Share:
TencentDB Agent Memory: From Flat Vector Piles to Traceable Layered Memory

The point of Agent memory is not to store more

When an AI Agent moves from one-turn questions to long-running tasks, memory quickly becomes a performance bottleneck. Search results, tool output, error traces, user preferences, and previous conversations all remain in context. The model appears to know more, but it pays a higher token cost and can lose the important signal among unrelated fragments.

TencentCloud’s open-source TencentDB Agent Memory takes a different approach: instead of flattening everything into a vector database, it separates memory into layers. The Agent reads a high-density structure first, then follows an identifier back to the original evidence only when necessary.

The project is positioned as a team memory hub for AI Agents. It turns conversations, documents, and code into reusable memory assets, and provides integrations for both OpenClaw and Hermes Agent. For anyone building multi-step Agents, tool-use workflows, or long-lived personal assistants, its value is not another chatbot. Its value is treating state management as a first-class part of Agent engineering.

The performance figures and architecture descriptions in this article come from the project README and official documentation. The benchmarks are reported by the authors and should not be treated as a guarantee for your own models or datasets.

Why traditional memory systems become unmanageable

The obvious implementation is to split each conversation into chunks, generate embeddings, and put them in a vector database. On the next request, the system retrieves a few similar passages. This works well for FAQs and short documents, but long-running Agents expose three structural problems.

First, vector similarity is not the same as workflow relationship. An error saying that a deployment failed and a preference saying that a user likes Docker may share technical vocabulary, but they play completely different roles in a task. Similarity alone does not tell the system which item is a fact, which is a scenario, and which is a reusable operating pattern.

Second, context has no cost awareness. A tool can return hundreds of thousands of tokens while the Agent only needs to know which step succeeded, which node failed, and what should happen next. Injecting the complete log wastes tokens and makes it easier for the model to overlook the important detail.

Third, summaries are often irreversible. If a summary omits a parameter or the cause of an error, it is difficult to reconstruct how the conclusion was produced, let alone return to the original tool output for verification.

TencentDB Agent Memory targets all three problems: layering instead of flat storage, symbolic state instead of verbose logs, and a traceable path from a high-level structure back to the original evidence.

Two pillars: layered memory and symbolic memory

1. Short-term memory: move tool logs out of context

Short-term memory answers the question: “What happened during this task?” The project uses three layers:

  • Bottom layer: Full tool output and raw text are written to external files such as refs/*.md, preserving the details needed for investigation.
  • Middle layer: Each step is extracted into structured summaries such as jsonl, recording the step, result, and relationships.
  • Top layer: A Mermaid graph compresses task state into a dense symbolic canvas. Only this lightweight structure is injected into the Agent context.

The model normally reads the top-level graph to understand the major task nodes. If it needs to inspect an error or a tool return value, it follows node_id to the raw text. This is progressive disclosure in practice: provide enough information for the current decision, and reveal detail on demand.

The conceptual data flow looks like this:

graph LR
    A[Full tool output] -->|Preserve raw text| B[refs/*.md]
    A -->|Extract relations| C[Mermaid symbolic canvas]
    C -->|Light injection| D[Agent context]
    D -.->|Recall by node_id| B

Unlike a plain summary, this strategy makes context compression and evidence preservation work together. The Agent does not carry the entire log on every turn, but the system does not throw away the source either.

2. Long-term memory: abstract conversations into a Persona

Cross-session memory answers a different question: “What patterns are stable for this person, team, or project?” TencentDB Agent Memory describes long-term personalization as a semantic pyramid:

  • L0 Conversation: The original dialogue.
  • L1 Atom: Atomic facts extracted from dialogue, such as preferences, constraints, or confirmed decisions.
  • L2 Scenario: A meaningful situation assembled from several facts.
  • L3 Persona: A more stable profile of a user or team.

These layers do not replace one another. They have different retrieval costs. An Agent can normally read a Persona or Scenario first, then drill down to an Atom or the original Conversation when a detail affects a decision. Remembering a preference therefore does not require loading the entire history every time.

The same idea extends to Skill generation. Repeated solutions are identified in execution traces, organized into Scenarios, and eventually distilled into reusable Skills or SOPs. For enterprise Agents, this is closer to knowledge operations than simply archiving chat history.

Traceability: why `node_id` matters

Memory systems have two dangerous failure modes: the model treats a wrong summary as truth, or an engineer cannot answer where a conclusion came from. The project preserves a drill-down chain for high-level symbols:

Persona / Mermaid Canvas
        ↓
Scenario / jsonl index
        ↓
Atom / refs
        ↓
Original conversations, tool output, and error traces

In short-term memory, nodes on the Mermaid canvas carry a node_id. The Agent uses the graph to reason about state transitions, while software uses the same ID to locate the raw file. This connects a model-readable summary with evidence that an engineer can inspect.

This does not eliminate hallucinations or guarantee that every extraction is correct. It does, however, create an actionable entry point for debugging. When an Agent says that deployment failed at step three because of a permission problem, the system should be able to take you to the relevant tool output instead of asking you to trust an ungrounded summary.

Getting started: begin with local SQLite

The official README provides OpenClaw and Hermes integration paths. The easiest way to validate the idea is to start locally with SQLite and sqlite-vec, rather than turning the first experiment into a remote database deployment.

OpenClaw: install and enable the plugin

openclaw plugins install @tencentdb-agent-memory/memory-tencentdb
openclaw gateway restart

Then enable the plugin in OpenClaw:

{
  "memory-tencentdb": {
    "enabled": true
  }
}

The default backend is local SQLite + sqlite-vec. After activation, the plugin handles conversation capture, memory extraction, scene aggregation, Persona generation, and recall before the next turn. To test short-term context offloading, add the offload setting and point the contextEngine slot at memory-tencentdb.

{
  "plugins": {
    "slots": {
      "contextEngine": "memory-tencentdb"
    }
  },
  "memory-tencentdb": {
    "enabled": true,
    "config": {
      "offload": {
        "enabled": true
      }
    }
  }
}

The README also notes that newer OpenClaw versions may require the after-tool-call-messages.patch.sh script so messages after tool calls can be offloaded and recovered correctly. Check your installed version and the official instructions before applying a patch.

Hermes: attach to an existing installation

If Hermes Agent is already installed, you can add the provider without using the dedicated Docker image. The core installation steps are:

mkdir -p ~/.memory-tencentdb
cd ~/.memory-tencentdb
npm init -y --silent
npm install @tencentdb-agent-memory/memory-tencentdb@latest --omit=dev

Place the plugin in a stable directory and link it to the Hermes provider directory:

rm -rf ~/.hermes/hermes-agent/plugins/memory/memory_tencentdb
ln -sf ~/.memory-tencentdb/tdai-memory-openclaw-plugin/hermes-plugin/memory/memory_tencentdb \\
  ~/.hermes/hermes-agent/plugins/memory/memory_tencentdb

There is an easy-to-miss naming detail: the directory must be named memory_tencentdb with an underscore. The configuration layer may use the memory-tencentdb alias, but the provider directory cannot use a hyphen.

Declare the provider in Hermes:

memory:
  provider: memory_tencentdb

The Gateway needs a start command and model configuration. Keep the API key in environment variables or a controlled .env file. Do not put credentials in an article, an Issue, or shell history:

MEMORY_TENCENTDB_GATEWAY_HOST="127.0.0.1"
MEMORY_TENCENTDB_GATEWAY_PORT="8420"
TDAI_LLM_BASE_URL="https://your-openai-compatible-endpoint/v1"
TDAI_LLM_MODEL="your-model"
TDAI_LLM_API_KEY="your-api-key"

The Gateway can be started manually, or the provider can discover and start it when the first conversation begins. Verify the service with its health endpoint:

curl http://127.0.0.1:8420/health

The response should contain a status of ok or degraded. With an existing Hermes installation, keep the original memory provider configuration and compare context length, latency, and error rate before and after enabling the new provider.

Docker: isolated testing or greenfield deployment

For a new Hermes instance with memory enabled, the project also documents a Docker path. The Gateway listens on 8420 and data is stored in a named volume. This makes dependencies easy to isolate and rebuild, but you still need to manage model credentials, volume backups, and network exposure.

docker build -f Dockerfile.hermes -t hermes-memory .
docker run -d \\
  --name hermes-memory \\
  --restart unless-stopped \\
  -p 8420:8420 \\
  -e MODEL_API_KEY="your-api-key" \\
  -v hermes_data:/opt/data \\
  hermes-memory
curl http://localhost:8420/health

The key in this example is only a placeholder. In production, use a secret manager, a restricted environment file, or the platform’s secret injection mechanism. Never commit a real key to Git.

How to read the project’s benchmark claims

The README reports several results after integrating with OpenClaw: on WideSearch, success rate rises from 33% to 50% while token usage falls from 221.31M to 85.64M; on SWE-bench, success rate rises from 58.4% to 64.2% while token usage falls from 3474.1M to 2375.4M; PersonaMem rises from 48% to 76%.

These numbers are interesting, but the correct interpretation is “improvement measured by the authors in a particular environment,” not “installation will always cut token usage by 61.38%.” Results depend on the model version, context strategy, number of tools, task length, data distribution, and extraction overhead.

To validate the system yourself, keep these variables fixed:

1. The same model and temperature.

2. The same task set and available tools.

3. The same context limit and retry policy.

4. Total tokens, per-turn latency, tool success rate, and final task completion rate.

5. Manual samples of memory extraction errors, not only the cost reduction.

The memory layer itself consumes model calls and storage. The useful metric is not token count in isolation, but the total cost of completing the same task set and whether results become more reproducible.

When is it a good fit?

Good fit:

  • Research, development, and automation Agents that execute many tool steps.
  • Workflows that repeatedly use the same SOPs, project context, and team conventions.
  • Teams that want governed memory shared across Agents or work sessions.
  • Systems that need lower context cost while retaining error investigation and evidence drill-down.
  • Existing OpenClaw or Hermes installations that want a provider/plugin instead of a complete Agent rewrite.

Not necessarily a fit:

  • A simple chatbot with only a few turns and no cross-session state.
  • One-off semantic search that does not need Scenarios, Personas, or tool-chain traceability.
  • Teams that cannot yet monitor retention, permissions, or extraction quality.

Any system that stores user preferences, conversations, and tool output must define data boundaries first. What may be stored long term? Who can query it? How is it deleted? How are sensitive fields handled? Layering solves retrieval and cost problems; it does not automatically solve governance.

A practical evaluation checklist

If you are considering the project for an existing Agent, test it in this order.

Step one: enable long-term memory only

Start with local SQLite and create repeatable scenarios such as output-format preferences, deployment rules, and common error fixes. Confirm that the next session retrieves the right information before adding short-term offloading.

Step two: add short-term compression

Choose a task with verbose tool output. Measure context size and final success rate. Check whether the Agent can find the correct node_id in the Mermaid canvas and retrieve the full source from refs when needed.

Step three: test failure and degradation

Temporarily make the Gateway unavailable. Observe whether the main Agent reports an error, skips memory, or blocks the whole task. Memory is an auxiliary capability; a database or Gateway failure should not take down the core conversation.

Step four: decide whether to centralize

Only after local tests pass should you evaluate team sharing, backups, permissions, and high availability. Back up raw evidence and high-level Markdown assets separately, and define deletion and retention procedures.

Conclusion: good Agent memory should be both cheaper and inspectable

The most interesting part of TencentDB Agent Memory is not the number of platforms it supports. It is the attempt to redefine Agent memory as a layered, compressible, traceable system. Short-term memory uses a symbolic canvas to reduce context load. Long-term memory uses Conversation → Atom → Scenario → Persona as abstraction layers. The bottom layer keeps evidence that can still be retrieved.

This design also offers a useful rule: an Agent remembering something should not mean pushing the entire history back into the prompt. Useful memory knows what stays resident, what can be loaded later, what must preserve the original text, and how a human can inspect the source when a result looks suspicious.

If you are building a multi-step Agent, a long-lived workflow, or a team AI assistant, this project is a strong reference implementation for a memory layer. The practical starting point is not to accept the benchmark blindly, but to build a small reproducible task set, measure tokens, latency, success rate, and traceability before deciding whether it belongs in production.

Further reading and sources