Skip to main content

Use project memory

Project memory keeps the conversations of your project and learns from them. Your agents and your apps then use what it learned.

Fig. 1 · How memory learns from your calls
Chat calls of your appskept when memory is on
Conversationsstored for search, no model call
Memoriesfacts, decisions, procedures, and a knowledge page
Used byyour agents over MCP or the API, and your apps with recall
You use memory fromWithTo
The consoleYour sign-inTurn memory on, review and pin memories, erase an end user
An MCP client, such as Claude CodeYour sign-inLet an agent read and save memories. See Connect an MCP client
Your codeA virtual keyRead, save and erase memories with the /v1/memory API
Each chat callA virtual keyAdd the matching memories to the call with a header

1. Memory is on for every project​

Memory is on for every project that has not turned it off. Memory needs every prompt and answer, so while it is on, Proxium stores every call, and you cannot narrow Stored prompts and answers.

To change memory:

  1. In the console, open Memory.
  2. Optional: set Memory to off. Proxium saves it at once.
  3. Optional: in Only these senders, type the apps to remember, separated by commas. Empty means every app.
  4. Optional: in Delete conversations after, type a number of days. Empty keeps every conversation and memory, which is the default.
  5. Optional: in Batch of and At least every, set when Proxium learns. See When Proxium learns. With a value in Delete conversations after, At least every must be shorter than it.
  6. Select Save.

Any member of the project can do this. An agent over MCP does the same with proxium_read_memory_settings and proxium_set_memory_settings, as the person who signed in. An admin of Proxium can change the memory of any project with PUT /provision/teams/{project}/memory. The Memory screen shows the Knowledge page, the Memories and the Recent conversations.

If you turn memory off, Stored prompts and answers stays at all. To store less, change What to store in Settings › Stored prompts and answers. How long they are kept is a separate setting there: see Data handling.

What Proxium remembers​

Proxium remembers a call when all of these are true:

ConditionExample
The route is /v1/chat/completions, /v1/messages or /v1/responsesNot /v1/embeddings
The call succeededA 503 upstream_failed is not remembered
The app is in Only these senders, or that field is emptyx-proxium-source: support-bot
The call does not send x-proxium-memory: off

A call without x-proxium-source is remembered only when Only these senders is empty.

When a conversation has no new turn for a while, Proxium reads it. It stores the text for search, and the first user message becomes the title. This step makes no model call.

Proxium reads the user and assistant messages, and the tool calls of the assistant: the name of the tool and its arguments. So a plan or a task that an agent creates with a tool goes into memory. Proxium does not read system prompts or tool results.

Which conversations Proxium learns from​

Every remembered conversation is stored for search. Proxium writes memories from a conversation only in these cases:

The conversationProxium writes memories
Has more than one turn: a person who chats, or an agent that works in stepsYes
Has an end user (a subject)Yes
Has a call with x-proxium-memory: writeYes
You asked Proxium to remember it. See Ask Proxium to remember a conversationYes
Is one call with no end user and no header, such as a job of a pipelineNo. It is stored for search only

A pipeline job, such as "draft a reply" or "score this text", is one call. Memories from it are seldom read, and each extraction is a model call. So Proxium skips it, unless the job sends x-proxium-memory: write.

When Proxium learns​

Fig. 2 · what memory does, and when
  1. Your app sends a call through Proxium to the vendor. Proxium stores the call in the conversation of the end user. This step makes no model call.
  2. When a conversation has no new turn for a while, Proxium reads it and stores the text for search. This step makes no model call.
  3. When enough new messages wait, or the oldest waited long enough, a batch starts. It makes one model call for each end user with new messages, and writes the memories. A memory that repeats a stored memory word for word is skipped with no model call.
  4. A later call that sends x-proxium-memory: recall gets the matching memories as one system message. This adds one embedding to the call.

Proxium finds the memories in batches, for each project. A batch starts when one of these is true:

ConditionField on the Memory screen
Enough new messages waitBatch of
The oldest new message waited long enoughAt least every

A batch makes one model call for each end user with new messages, and one for the calls with no end user. Each call reads only the messages since the last batch. The cost follows the amount of text that Proxium learns from, not the number of batches.

The batch uses the models of the tier memory-extract. The default is glm-5.3-flash at reasoning low, then gemma4:31b. To use your own models, set the tier memory-extract on Routing.

The Memory screen shows how many new messages wait, and the latest time of the next batch. If a batch fails, the screen shows the error. Proxium tries again later, with longer waits, and the messages wait until a batch succeeds.

2. Choose what each call does​

Send the x-proxium-memory header on a call:

ValueWhat Proxium does
No headerRemembers the call. Proxium learns from it by the rules in Which conversations Proxium learns from
writeRemembers the call, and Proxium learns from its conversation, even when it is one call with no end user
recallRemembers the call, and first adds the matching memories to it. Only on chat and Messages calls
offDoes not remember the call

Example: a support bot that knows the earlier tickets of the user.

Terminal
curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-H "x-proxium-source: support-bot" \
-H "x-proxium-memory: recall" \
-H "x-proxium-memory-subject: user-ana" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"messages\": [{\"role\": \"user\", \"content\": \"Which plan am I on?\"}]}"

With recall, Proxium searches the memories for the last user message. It adds the best matches as one system message, after your own system messages. If the search is slow or fails, the call continues without memories.

On /v1/responses, recall acts as write. Any other value gets 400 invalid_memory_mode.

3. Name the end user​

A memory can be about one end user of your app. Proxium calls that end user the subject. It reads the subject from:

  1. The x-proxium-memory-subject header.
  2. If there is no header, the user field of the body. On /v1/messages, metadata.user_id.
CallReads the memories of
With subject user-anauser-ana, and the memories of the whole project
Without a subjectThe whole project, and the memories of every end user
warning

Send the subject on each recall call that is about one end user. Without it, the call can get the memories of other end users.

Use an opaque id, not an email address. A subject has at most 256 characters, and no control characters. Else a remembered call gets 400 invalid_memory_subject.

4. Give your agents the memory​

An agent with an MCP client uses the MCP server. Your own code uses the memory API with a virtual key, at https://proxium.tech/v1:

TaskRequestReference
Read the memory at the start of a taskPOST /v1/memory/wake-upWake-up
Search conversations and memoriesPOST /v1/memory/searchSearch
Read one conversationGET /v1/memory/episodes/{id}Episode
Save a memoryPOST /v1/memory/itemsRemember
Ask Proxium to remember a conversationPOST /v1/memory/episodes/{id}/rememberRemember a conversation
Correct a memoryPATCH /v1/memory/items/{id}Correct
Rate a memoryPOST /v1/memory/items/{id}/ratingRate
Remove a memoryDELETE /v1/memory/items/{id}Forget
Read the knowledge graphPOST /v1/memory/graphGraph

Start a task with wake-up​

The answer holds the pinned and trusted memories, the titles of the newest conversations and the knowledge page. A subject adds the memories of that end user.

Terminal
curl https://proxium.tech/v1/memory/wake-up \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"subject": "user-ana"}'

Save a memory​

Save one statement in each memory. kind is fact (default), preference, decision, event or procedure. importance says how much a memory counts. Leave out subject for a memory of the whole project.

Terminal
curl https://proxium.tech/v1/memory/items \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"text": "Large refunds need a second approval.", "kind": "procedure"}'

Ask Proxium to remember a conversation​

Save a memory when you know the statement. When a whole conversation matters, ask Proxium to read it and write the memories. Proxium uses its memory model, and the batch starts at once. Get the id of the conversation from search or wake-up. Proxium reads the conversation from its start. focus is optional: it says what to look for.

Terminal
curl https://proxium.tech/v1/memory/episodes/$CONVERSATION_ID/remember \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"focus": "the plan and the tasks"}'

Proxium answers 202. An agent over MCP uses the tool proxium_remember_conversation.

Rate a memory after you use it​

Terminal
curl https://proxium.tech/v1/memory/items/$MEMORY_ID/rating \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"rating": "useful"}'

A memory rated wrong two times leaves the recall, until a person selects Confirm on the Memory screen.

What a virtual key may change​

ActionAllowedIf not
Save a memoryWhen memory is on409 memory_off
Pin a memoryNever. Only a person can pin, in the console403 forbidden
Correct or remove a memory of an agentYes
Correct or remove a memory that a person wroteNo403 forbidden

A text has a size limit. Proxium refuses a text that looks like a prompt injection (403), and it replaces a credential in the text with a placeholder. A correction keeps the old memory in the history.

Review what agents saved​

Open the Memory screen. The Memories list shows to review first: the memories of agents that no person confirmed. For each one, select Confirm, Pin, Edit or Remove. To add a memory yourself, fill in New memory, Kind and User, and select Save.

Erase one end user​

warning

You cannot undo an erasure.

With the API, with any virtual key of the project
Terminal
curl -X DELETE https://proxium.tech/v1/memory/subjects/user-ana \
-H "Authorization: Bearer $PROXIUM_KEY"
In the console, as an owner of the project
  1. Open Memory. The section Erase one user is at the bottom of the screen.
  2. Type the subject in User id.
  3. Select Erase permanently, and confirm.

Proxium deletes the memories, the profile and the conversations of the subject, in one transaction. It also deletes the stored prompts and answers of every call that named that subject, whether memory kept the call or not. The audit record keeps a SHA-256 hash of the subject. Data handling says what stays.

Delete old conversations​

By default, Proxium keeps every conversation and memory. To delete old conversations, type a number of days in Delete conversations after on the Memory screen. Once a day, Proxium then deletes the conversations whose last turn is older than that. It also deletes the memories that were corrected or removed before that time. Live memories stay. A conversation that memory has not read yet stays until memory reads it. The deletion runs with memory on or off.

What memory costs​

The memories, the pages and the embeddings come from model calls. Reading a conversation makes none. Proxium sends them through the routing of your project, and records them as your spend, with the application proxium-memory. The top of the Memory screen shows the memory cost of this month.

Proxium uses the first tier of each row that has a chain in your project:

JobTiers, in order
Memories from the conversations, in batchesmemory-extract, then trivial
Merges, graph facts and pagesmemory-consolidate, then memory-extract, then trivial
Embeddingsmemory-embed, then embed

To set a chain for a tier, see Route requests to models.