Conversation
This was referenced Sep 21, 2026
Merged
Merged
This was referenced Sep 21, 2026
Draft
[session storage 2/2] Remove open-or-create from both stores, and close from the segment store
#1625
Draft
[qdrant options] Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization
#1618
Draft
edwinyyyu
force-pushed
the
feat/vector-store-declared-routing-main
branch
from
September 21, 2026 22:44
48951bf to
c164c56
Compare
edwinyyyu
force-pushed
the
feat/vector-store-declared-routing-main
branch
from
September 25, 2026 23:10
c164c56 to
2dc8c96
Compare
When a session's collection stayed pending, or kept losing races, the retry loop gave up with a RuntimeError that carried no cause, so the log lost the pending error's registered_at, the time an operator needs to tell an abandoned creation from a slow one. The RuntimeError is now raised from the last error the loop caught; its type and message are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The claim computes 1 << (failed_rounds - 1) for each candidate row, rows at 0 failures included, where the shift count is -1: defined on SQLite, undefined behavior inside PostgreSQL's int4shl. The row's first disjunct made it claimable regardless, so no claim went wrong, but the expression was undefined there. The count is now clamped at 0 with a CASE, which both dialects evaluate alike; backoffs after a failure are unchanged, and SQLite's plan for the claim is the same index range on (vector_store_name, enqueued_at). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When every attempt lost the reservation or the confirmation to another creator or a deleter, open_or_create_collection gave up with a VectorStoreAttemptsExhaustedError that carried no cause. It is now raised from the last lost race, as the event backend's locator chains its own give-up error; its type and message are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A query whose limit is not positive is refused with the other invalid inputs, before the handle checks its liveness, so it is no longer an example of a call with nothing to send that still checks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The error log fired only when a failed round brought the count to exactly the bound. The claim takes only tombstones under the bound, so a tombstone is dead-lettered once its count reaches the bound or goes past it, as when two purgers fail one tombstone at once; the log now fires in either case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
claim_purgeable_incarnation() was an async context manager, which cannot see what its body returns, so a round reported whether it found records through PurgeClaim.any_records_found: a field that started None and a guard that raised when a round left it unset. run_purge_round() takes the round as a callable instead. The registry claims the tombstone that came due first, calls the round with its namespace, configuration and incarnation, and records what the round returns under the claim. A round that raises still counts as a failed round. PurgeClaim, its None state and the guard go; the vector store's _purge_round hook is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On SQLite the claim was a SELECT whose FOR UPDATE SKIP LOCKED the dialect drops, and the driver defers BEGIN until a write, so the claim ran outside any transaction: two processes sharing a SQLite registry could claim one tombstone at once. The round recorded afterwards then decided the failed-round reset from what it read at claim time, so a round that succeeded could skip the reset after a racing round failed. On SQLite the claim is now the segment store's: an UPDATE ... RETURNING on the oldest eligible tombstone, which opens the write transaction, so purgers serialize at the claim and rounds run one at a time. PostgreSQL keeps FOR UPDATE SKIP LOCKED. The cost on SQLite is that its single write lock is held across the round's remote deletion: the registry's other writers wait for it, and past the driver's busy timeout they fail with a locked-database error. The purge document states it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A creation whose storage preparation failed or was cancelled starts the reservation's cancel as a shielded task and logged its failure around the await. A creation cancelled again stops awaiting it, so a cancel that then failed went unlogged in the store, surfacing only as asyncio's "Task exception was never retrieved" without the collection's name. The task's done-callback now reports the failure, whoever is still awaiting it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
utils.require_identifiers checked the namespace and name of the registry-backed base's lifecycle calls, while the four stores kept the check inline: Qdrant and Milvus with the helper's own body, the SQLite stores with one combined check whose message, "Invalid namespace ... or name ...", named neither the rule nor which identifier broke it. Every store's lifecycle calls now use the helper, so an invalid identifier is reported the same way everywhere, with the rule it breaks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The registry-backed store's create, open-or-create, delete and purge ran under the store's operation tracker, and open_collection, which resolves the name in the registry and is the first call the event backend's locator makes on every open, did not. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The deployment note said existing Qdrant and Milvus data is orphaned. The native collections keep their names, so the upgrade's new data lands beside the old: a Qdrant collection kept through it holds old points no search sees and no purge reclaims, and dropping it afterwards drops the new data too; an existing Milvus collection has the earlier schema, which the store cannot prepare. The note now says to drop the collections before upgrading, and why. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A creation cancelled its reservation when preparing the collection's storage raised or was cancelled, so the name was free for the next attempt, but not when the confirmation that follows did: a failure or a cancellation while confirming left the collection pending, its name taken until someone deleted it. Preparation and confirmation now share one cleanup, the same shielded, logged cancel. The cancel acts only on a pending collection, so a confirmation that committed before its failure or cancellation was observed stands. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The purge document said the registry's other writers wait while a SQLite round holds the write lock. The lock covers the whole database file, so every store writing to it waits, and under the configuration wizard's defaults the episode store, session manager, segment store and configuration database share the registry's file. The round holds the lock across remote calls bounded by the store's request timeout, and past SQLite's default 5 s busy timeout, which the server does not change, those writers fail with a locked-database error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The integration tests ran Qdrant 1.17.0. The Qdrant store's design, which follows in this stack, was measured against 1.19.1: its one-shot purge of an incarnation by filter, and the per-tenant index layout. 1.17.0 predates the filter-resolution fence that 1.19.0 added to filter deletes (qdrant#9678), whose extra cost 1.19.1 no longer shows, so the tests could not see the behavior the store is tuned for. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Qdrant client was built with qdrant-client's default timeout, so how long a write to Qdrant can be in flight was nothing the configuration stated. The purge that follows in this stack waits out a retention longer than any write can be in flight, and the request timeout is the part of that time the store controls. `QdrantConf.request_timeout_seconds`, a positive whole number of seconds defaulting to 30, is passed to the client. The sample configurations and the configuration docs show it. The tests' Qdrant clients are built by one fixture with the configuration's default, so they run with the timeout a configured store has. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Qdrant store kept its catalog in a `__registry` collection per namespace and serialized creating and deleting a collection with locks in one process. A collection's name was the tenant discriminator on its points, so a handle kept writing into a collection deleted and created again under its name (MemMachine#1563), and a write in flight during a deletion outlived it. `QdrantVectorStore` is now a `RegistryBackedVectorStore`, and any number of processes sharing its registry may serve its collections: - Its catalog is the collection registry in the relational database `QdrantConf.collection_registry` names. The `__registry` collections, `registry_replication_factor` and the process-local locks go. The database manager builds and starts the registry before it opens the client, so a registry database it cannot resolve leaves no client open. - Every point carries its collection's incarnation in `sys-incarnation`, in place of the name, under the same tenant index, and every search filters on it. - A point's id is `uuid5(incarnation, record UUID)`, and the record UUID is kept in the payload as `sys-record_uuid`, which a search returns. The logical collections sharing a native collection share its id space: with the record UUID as the id, an upsert of a UUID another collection held replaced that collection's point. - Preparing a collection's storage creates its native collection and payload indexes, each under its own already-exists guard, so a creation that failed part way is completed by the next. - A purge round looks for one point under the incarnation and, finding one, deletes the incarnation's points with one filter-delete. `QdrantConf.tombstone_retention_seconds`, a day by default, is the retention, and the configuration refuses one below 10 x `request_timeout_seconds` + 300 seconds. - An upsert is halved only when Qdrant or a proxy refuses it as sent, with a 400 or a 413. Any other error raises at once: a timed-out upsert may still be applied, and sending it again adds load to a server already too slow. The collection lifecycle contract (`collection_lifecycle_contract.py`), which a store's tests mix in with hooks that read the backend directly, runs on Qdrant: stale handles, a collection created again starting empty, creation races and their outcomes, failed preparations, the purge, a write landing under a dead incarnation, and the registry lookups each operation makes. The store's own tests check what Qdrant holds by scrolling the incarnation past the store. Breaking: existing Qdrant data is orphaned, since its points carry names and its catalog is in the `__registry` collections; no migration is included. Every Qdrant store needs a relational database for its registry, which the samples, the Helm chart, the configuration wizard and the configuration docs name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`design/qdrant_vector_store.md` records the Qdrant store's layout, its derived point ids and why they are one-way, filtered-search correctness, the purge by filter, and its consistency on one node and replicated, with the measurements behind each choice. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The collection registry now answers registrations: a PendingRegistration from register, which its creator marks live, and a LiveRegistration from resolve and mark_live, whose require_current is a handle's fence. The Qdrant handle takes its live registration in place of the namespace, name, incarnation, configuration and registry lookup, and _build_collection_handle builds it from one. The collection lifecycle contract follows: - the tests that fail a check or count checks patch the registration type's require_current, where they replaced the handle's lookup; - the racing winners register and mark live through a pending registration, and the winner is found with resolve; - churn counts VectorStoreCollectionDeletedError, which a creation undone by a concurrent deletion now raises, among the domain's outcomes. The Qdrant tests that build a handle on a mocked client give it a live registration that stays current. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… and the contract MemMachine#1734 names the registry's handles for their holders: `reserve` answers a Reservation, whose `confirm` answers a Registration, and whose `cancel` gives the name back. The Qdrant handle takes a Registration, and the collection lifecycle contract's racing winners reserve, then confirm. Names only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A limit at or below zero is now refused as invalid input, checked before the handle's liveness, so it no longer stands for a query with nothing to send to the backend; the query with no vectors still does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…istry The user docs did not say what happens to existing Qdrant data. The store keeps its native collections' names, so points left in one stay, invisible and never purged, and dropping the collection after the upgrade drops the new data too; databases.mdx now says to drop the collections before upgrading. The Helm README's configuration summary also lists collection_registry, which the configmap template sets. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…server
The Milvus store kept its catalog in a `memmachine_<namespace>__registry`
collection per namespace, whose `insert` does not enforce primary-key
uniqueness, so two creators of one name both succeeded, and it serialized
creating and deleting a collection with locks in one process. A
collection's name was the partition key on its entities, so a handle kept
writing into a collection deleted and created again under its name, and a
write in flight during a deletion outlived it. Each call ran the sync
client on a thread of the event loop's default executor, so calls waiting
on Milvus queued every other `to_thread` call of the process behind them.
`MilvusVectorStore` is now a `RegistryBackedVectorStore`, and any number
of processes sharing its registry may serve its collections. The store
was rewritten rather than adapted; what changes:
- The catalog is the collection registry in the relational database
`MilvusConf.collection_registry` names, built and started by the
database manager before it opens the client. The registry collections
and the process-local locks go.
- Every entity carries its collection's incarnation as the partition key,
with partition-key isolation, and in its primary key,
`"{incarnation}:{record_uuid}"`.
- Declared properties are typed, indexed fields (`_p_<name>`, a datetime
as TIMESTAMPTZ with its UTC offset beside it in `_tz_<name>`), and
undeclared ones live in the JSON field, still filterable. Negation is
the complement, as on Qdrant: a negated condition holds where the
property has no value.
- The vector index is HNSW_SQ with 4-bit codes and FP16 refinement, what
AUTOINDEX builds on CPU from Milvus 2.6.10, named so every server builds
the same; a search rescores `limit x 8` candidates. Scores are the
server's.
- Reads run at Milvus's default consistency level, Bounded, which the
store states: a query reflects every write made at least the server's
`common.gracefulTime` before it. `MilvusConf.consistency_level` goes.
- The store calls Milvus through pymilvus's `AsyncMilvusClient`, and
bounds every request by `MilvusConf.request_timeout_seconds`.
- A purge round lists a batch of the incarnation's primary keys, at most
`MilvusConf.purge_batch_size`, and deletes them.
`MilvusConf.tombstone_retention_seconds` is the retention, refused below
10 x `request_timeout_seconds` + 300 seconds as for Qdrant. A declared
string's VARCHAR length is `MilvusConf.max_varchar_length`. Limits the
server configures stay the server's.
- A delete raises unless Milvus accepted every key sent.
Milvus Lite is dropped. It is a separate embedded engine that scores,
indexes and enforces collection properties differently, so a store tested
against it is not tested against what production runs; MemMachine's
local, single-node backend is the SQLite vector store. `MilvusConf.uri`
defaults to `http://localhost:19530`, a URI with no scheme, which pymilvus
reads as a Lite file, is refused, and the milvus extra no longer installs
milvus-lite (the lock drops it with the packages only it required).
The store's tests, including the collection lifecycle contract, run
against a Milvus 2.6.24 server container as integration tests, and read
past the store at Strong to check what Milvus holds.
Breaking: existing Milvus data is orphaned, and an existing native Milvus
collection has to be dropped; no migration is included. Every Milvus
store needs a relational database for its registry, which the samples,
the configuration wizard and the configuration docs name.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`design/milvus_vector_store.md` records the Milvus store's layout: the shared native collection and partition-key tenancy, the composite key, the index and why it was chosen, declared properties as typed fields, the purge in batches, consistency levels, and the async client, with the measurements behind each choice. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
As for Qdrant: the Milvus handle takes its live registration in place of the namespace, name, incarnation, configuration and registry lookup, and _build_collection_handle builds it from one. The test that builds a handle on a mocked client gives it a live registration that stays current. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Milvus store keeps the native collection's name from main, so on an upgraded server it finds main's collection, whose schema it cannot prepare, and the first request of every new session fails. The user docs' upgrade note now names Milvus beside Qdrant. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
edwinyyyu
force-pushed
the
feat/vector-store-declared-routing-main
branch
from
October 3, 2026 03:00
21d9145 to
a468529
Compare
…ck) (MemMachine#1606) * Regenerate the OpenAPI document under the locked FastAPI `docs/openapi.json` predates the FastAPI release in `uv.lock` (0.141.1), whose `ValidationError` component carries `input` and `ctx`; regenerating the document with `docs/tools/generate_openapi.py` adds the two fields and changes nothing else. Separate from the API changes above it so their diffs of this file show only what they change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn * Remove per-project filterable properties A project could declare `properties_schema`, a set of caller property keys with types, on its long-term memory configuration; the event backend merged it into the vector store collection's indexed schema and rejected filters on any other `m.<key>`. That let a tenant create database resources (indexes, columns) by naming them in a request, which is what forced per-collection native resources named by a hash of their schema on the backends that limit them. The option is removed from the server configuration, the project API and the memory-configuration API, the Python SDK, the sample configurations, the configuration docs and the OpenAPI document. A filter may name any `m.<key>`; the stores evaluate it on the properties they hold. What a store indexes is decided by the deployment, not per project. A breaking API change on `speedkick`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Rebased onto MemMachine#1631: the per-project schema also leaves MemMachine#1631's service locator, which creates the session's collection in a retry loop, and the commented option goes from the event sample configuration MemMachine#1698 added. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit a8322a7)
The event backend wrote every property of an event into its vector record and mapped the caller's whole filter onto the vector store, so a user key that a deployment never declared was both stored and filtered there: on Qdrant and Milvus as unindexed payload a filtered query scans for. The segment store already holds every property and already receives the whole filter for the context windows, so the vector side only duplicated work the segment store does anyway. The vector record now carries the keys the collection declares: EventMemory's reserved timestamp, and the `_`-prefixed system properties an adapter stamps on the event, which after MemMachine#1670 are exactly the collection's schema. The vector store is queried with the conjuncts of the filter that name only such fields; a conjunct is dropped whole when any field under it is a user property, so dropping only ever widens the vector search, and the segment store narrows it back on the windows. User keys never reach the vector store, so a tenant's properties cannot shape what it stores or scans. `filter_fields` joins the filter parser: every field name a tree addresses. Rebased onto MemMachine#1631, where the vector record no longer carries the segment uuid (the segment store maps a derivative to its segment). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu
force-pushed
the
feat/vector-store-declared-routing-main
branch
from
October 3, 2026 04:00
a468529 to
2376135
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose of the change
The event backend wrote every property of an event into its vector record and mapped the caller's whole filter onto the vector store, so a user key a deployment never declared was both stored and filtered there: on Qdrant and Milvus as unindexed payload a filtered query scans for. With #1670 a tenant can no longer declare indexes; this closes the other half, so a tenant's properties cannot shape what the vector store stores or scans either. The segment store already holds every property and already receives the whole filter for the context windows, so the vector side only duplicated work the segment store does anyway.
The vector record now carries the keys the collection declares: EventMemory's reserved timestamp, and the
_-prefixed system properties an adapter stamps on the event, which after #1670 are exactly the collection's schema (collection.config.indexed_properties_schema). The vector store is queried with the conjuncts of the filter that name only such fields; a conjunct is dropped whole when any field under it is a user property, so dropping only ever widens the vector search, and the segment store narrows it back on the windows. User keys never reach the vector store.filter_fieldsjoins the filter parser: every field name a tree addresses. Tests: a user property stays on the segment and off the record; a user-property conjunct reaches only the segment store; a user property under anORleaves the vector search unfiltered;filter_fieldsnames every field under every node.This is the routing half of what was #1628 on
speedkick(its store-side half, the stores rejecting undeclared keys, is deferred to after #1627); it is re-derived onto #1736's collection shape from the same commit, where the segment store, not the vector record, maps a derivative to its segment.Stack
21 open PRs: one independent PR, and the vector store tree of short parallel branches. Every PR's GitHub base is
main, since the branches are in a fork and a pull request can target only this repository's branches; the on column gives the order the PRs build on each other instead. A stacked PR's diff on GitHub includes the PRs under it until they merge.Independent of the vector store tree, directly on
main:mainThe vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1670 and #1702 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.
mainmainThis PR is its one commit,
23761359f, stacked on #1670. #1627 is stacked on it.Verification
At every commit from #1736's head
b3baee60eto this PR's head23761359f(this PR and the PRs under it), on 2026-10-02:ruff checkandruff format --checkclean;ty checkclean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); the server suite without integration tests passes, 2104 tests at this PR's head. The integration tests of the vector stores, the resource manager, episodic memory and semantic storage, against PostgreSQL 16, Qdrant 1.19.1 and Milvus 2.6.24 in containers, pass at the heads of #1627 (603), #1625 (575), #1676 (600), #1618 (600) and #1616 (670).🤖 Generated with Claude Code
https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn