Use project memory
Project memory keeps the conversations of your project and learns from them. Your agents and your apps then use what it learned.
| You use memory from | With | To |
|---|---|---|
| The console | Your sign-in | Turn memory on, review and pin memories, erase an end user |
| An MCP client, such as Claude Code | Your sign-in | Let an agent read and save memories. See Connect an MCP client |
| Your code | A virtual key | Read, save and erase memories with the /v1/memory API |
| Each chat call | A virtual key | Add the matching memories to the call with a header |
1. Memory is on for every project
Memory is on for every project that has not turned it off. Memory needs every prompt and answer, so while it is on, Proxium stores every call, and you cannot narrow Stored prompts and answers.
To change memory:
- In the console, open Memory.
- Optional: set Memory to off. Proxium saves it at once.
- Optional: in Only these senders, type the apps to remember, separated by commas. Empty means every app.
- Optional: in Delete conversations after, type a number of days. Empty keeps every conversation and memory, which is the default.
- Optional: in Batch of and At least every, set when Proxium learns. See When Proxium learns. With a value in Delete conversations after, At least every must be shorter than it.
- Select Save.
Any member of the project can do this. An agent over MCP does the same with proxium_read_memory_settings and proxium_set_memory_settings, as the person who signed in. An admin of Proxium can change the memory of any project with PUT /provision/teams/{project}/memory. The Memory screen shows the Knowledge page, the Memories and the Recent conversations.
If you turn memory off, Stored prompts and answers stays at all. To store less, change What to store in Settings › Stored prompts and answers. How long they are kept is a separate setting there: see Data handling.
What Proxium remembers
Proxium remembers a call when all of these are true:
| Condition | Example |
|---|---|
The route is /v1/chat/completions, /v1/messages or /v1/responses | Not /v1/embeddings |
| The call succeeded | A 503 upstream_failed is not remembered |
| The app is in Only these senders, or that field is empty | x-proxium-source: support-bot |
The call does not send x-proxium-memory: off |
A call without x-proxium-source is remembered only when Only these senders is empty.
When a conversation has no new turn for a while, Proxium reads it. It stores the text for search, and the first user message becomes the title. This step makes no model call.
Proxium reads the user and assistant messages, and the tool calls of the assistant: the name of the tool and its arguments. So a plan or a task that an agent creates with a tool goes into memory. Proxium does not read system prompts or tool results.
Which conversations Proxium learns from
Every remembered conversation is stored for search. Proxium writes memories from a conversation only in these cases:
| The conversation | Proxium writes memories |
|---|---|
| Has more than one turn: a person who chats, or an agent that works in steps | Yes |
| Has an end user (a subject) | Yes |
Has a call with x-proxium-memory: write | Yes |
| You asked Proxium to remember it. See Ask Proxium to remember a conversation | Yes |
| Is one call with no end user and no header, such as a job of a pipeline | No. It is stored for search only |
A pipeline job, such as "draft a reply" or "score this text", is one call. Memories from it are seldom read, and each extraction is a model call. So Proxium skips it, unless the job sends x-proxium-memory: write.
When Proxium learns
- Your app sends a call through Proxium to the vendor. Proxium stores the call in the conversation of the end user. This step makes no model call.
- When a conversation has no new turn for a while, Proxium reads it and stores the text for search. This step makes no model call.
- When enough new messages wait, or the oldest waited long enough, a batch starts. It makes one model call for each end user with new messages, and writes the memories. A memory that repeats a stored memory word for word is skipped with no model call.
- A later call that sends
x-proxium-memory: recallgets the matching memories as one system message. This adds one embedding to the call.
Proxium finds the memories in batches, for each project. A batch starts when one of these is true:
| Condition | Field on the Memory screen |
|---|---|
| Enough new messages wait | Batch of |
| The oldest new message waited long enough | At least every |
A batch makes one model call for each end user with new messages, and one for the calls with no end user. Each call reads only the messages since the last batch. The cost follows the amount of text that Proxium learns from, not the number of batches.
The batch uses the models of the tier memory-extract. The default is glm-5.3-flash at reasoning low, then gemma4:31b. To use your own models, set the tier memory-extract on Routing.
The Memory screen shows how many new messages wait, and the latest time of the next batch. If a batch fails, the screen shows the error. Proxium tries again later, with longer waits, and the messages wait until a batch succeeds.
2. Choose what each call does
Send the x-proxium-memory header on a call:
| Value | What Proxium does |
|---|---|
| No header | Remembers the call. Proxium learns from it by the rules in Which conversations Proxium learns from |
write | Remembers the call, and Proxium learns from its conversation, even when it is one call with no end user |
recall | Remembers the call, and first adds the matching memories to it. Only on chat and Messages calls |
off | Does not remember the call |
Example: a support bot that knows the earlier tickets of the user.
curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-H "x-proxium-source: support-bot" \
-H "x-proxium-memory: recall" \
-H "x-proxium-memory-subject: user-ana" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"messages\": [{\"role\": \"user\", \"content\": \"Which plan am I on?\"}]}"
With recall, Proxium searches the memories for the last user message. It adds the best matches as one system message, after your own system messages. If the search is slow or fails, the call continues without memories.
On /v1/responses, recall acts as write. Any other value gets 400 invalid_memory_mode.
3. Name the end user
A memory can be about one end user of your app. Proxium calls that end user the subject. It reads the subject from:
- The
x-proxium-memory-subjectheader. - If there is no header, the
userfield of the body. On/v1/messages,metadata.user_id.
| Call | Reads the memories of |
|---|---|
With subject user-ana | user-ana, and the memories of the whole project |
| Without a subject | The whole project, and the memories of every end user |
Send the subject on each recall call that is about one end user. Without it, the call can get the memories of other end users.
Use an opaque id, not an email address. A subject has at most 256 characters, and no control characters. Else a remembered call gets 400 invalid_memory_subject.
4. Give your agents the memory
An agent with an MCP client uses the MCP server. Your own code uses the memory API with a virtual key, at https://proxium.tech/v1:
| Task | Request | Reference |
|---|---|---|
| Read the memory at the start of a task | POST /v1/memory/wake-up | Wake-up |
| Search conversations and memories | POST /v1/memory/search | Search |
| Read one conversation | GET /v1/memory/episodes/{id} | Episode |
| Save a memory | POST /v1/memory/items | Remember |
| Ask Proxium to remember a conversation | POST /v1/memory/episodes/{id}/remember | Remember a conversation |
| Correct a memory | PATCH /v1/memory/items/{id} | Correct |
| Rate a memory | POST /v1/memory/items/{id}/rating | Rate |
| Remove a memory | DELETE /v1/memory/items/{id} | Forget |
| Read the knowledge graph | POST /v1/memory/graph | Graph |
Start a task with wake-up
The answer holds the pinned and trusted memories, the titles of the newest conversations and the knowledge page. A subject adds the memories of that end user.
curl https://proxium.tech/v1/memory/wake-up \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"subject": "user-ana"}'
Save a memory
Save one statement in each memory. kind is fact (default), preference, decision, event or procedure. importance says how much a memory counts. Leave out subject for a memory of the whole project.
curl https://proxium.tech/v1/memory/items \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"text": "Large refunds need a second approval.", "kind": "procedure"}'
Ask Proxium to remember a conversation
Save a memory when you know the statement. When a whole conversation matters, ask Proxium to read it and write the memories. Proxium uses its memory model, and the batch starts at once. Get the id of the conversation from search or wake-up. Proxium reads the conversation from its start. focus is optional: it says what to look for.
curl https://proxium.tech/v1/memory/episodes/$CONVERSATION_ID/remember \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"focus": "the plan and the tasks"}'
Proxium answers 202. An agent over MCP uses the tool proxium_remember_conversation.
Rate a memory after you use it
curl https://proxium.tech/v1/memory/items/$MEMORY_ID/rating \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"rating": "useful"}'
A memory rated wrong two times leaves the recall, until a person selects Confirm on the Memory screen.
What a virtual key may change
| Action | Allowed | If not |
|---|---|---|
| Save a memory | When memory is on | 409 memory_off |
| Pin a memory | Never. Only a person can pin, in the console | 403 forbidden |
| Correct or remove a memory of an agent | Yes | |
| Correct or remove a memory that a person wrote | No | 403 forbidden |
A text has a size limit. Proxium refuses a text that looks like a prompt injection (403), and it replaces a credential in the text with a placeholder. A correction keeps the old memory in the history.
Review what agents saved
Open the Memory screen. The Memories list shows to review first: the memories of agents that no person confirmed. For each one, select Confirm, Pin, Edit or Remove. To add a memory yourself, fill in New memory, Kind and User, and select Save.
Erase one end user
You cannot undo an erasure.
With the API, with any virtual key of the project
curl -X DELETE https://proxium.tech/v1/memory/subjects/user-ana \
-H "Authorization: Bearer $PROXIUM_KEY"
In the console, as an owner of the project
- Open Memory. The section Erase one user is at the bottom of the screen.
- Type the subject in User id.
- Select Erase permanently, and confirm.
Proxium deletes the memories, the profile and the conversations of the subject, in one transaction. It also deletes the stored prompts and answers of every call that named that subject, whether memory kept the call or not. The audit record keeps a SHA-256 hash of the subject. Data handling says what stays.
Delete old conversations
By default, Proxium keeps every conversation and memory. To delete old conversations, type a number of days in Delete conversations after on the Memory screen. Once a day, Proxium then deletes the conversations whose last turn is older than that. It also deletes the memories that were corrected or removed before that time. Live memories stay. A conversation that memory has not read yet stays until memory reads it. The deletion runs with memory on or off.
What memory costs
The memories, the pages and the embeddings come from model calls. Reading a conversation makes none. Proxium sends them through the routing of your project, and records them as your spend, with the application proxium-memory. The top of the Memory screen shows the memory cost of this month.
Proxium uses the first tier of each row that has a chain in your project:
| Job | Tiers, in order |
|---|---|
| Memories from the conversations, in batches | memory-extract, then trivial |
| Merges, graph facts and pages | memory-consolidate, then memory-extract, then trivial |
| Embeddings | memory-embed, then embed |
To set a chain for a tier, see Route requests to models.