QdrantVFS exposes a Qdrant collection as a read-only filesystem: group-by
payload fields become nested folders, each point is a .json payload file (plus
a .txt text file and an optional blob), and semantic search is the search
command. The TypeScript backend mirrors the Python one
and returns identical results.
Install
The client is browser-safe, so the VFS ships incore and is available
from both the Node and browser packages:
Document and chunk lineage
Config fields accept Qdrant’s dotted nested-key syntax. A LangChain-style payload withpage_content and metadata: { source, page } can therefore use:
s3://docs/policies/refund-2026.pdf and page value 004,
the chunk text is exposed as
refund-2026.pdf/004__<point-id>.txt. The point id remains as a stable suffix,
so duplicate labels cannot collide and direct reads work without a warm cache.
Only the stem the listing publishes opens: another label in front of the same
id reads as absent. A label that is not a string spells as compact JSON
(true, 1, 1e-7), the same in TypeScript and Python. basenameFields strips
URL/path parents from the named groupBy fields; omit a field to preserve its
complete value in one path-safe segment: / renders as ∕, and a blank or
dot-led value is led by ⁄, so every value has its own segment that lists and
opens. A basename longer than 255 bytes, which ext4
and APFS refuse, is cut to fit and ends in __ plus the md5 of the whole name,
so two long leaves stay two directories. Basenames must be unique within a
parent group; ambiguous listings are refused, and opening a basename directory
checks every point of its parent group, so a second source past maxRows is
refused rather than hidden.
Filesystem layout
<id> is the Qdrant point id. With nameField, leaf stems use
<name>__<id>. Semantic search is the search command, not a path: it returns
ranked points as the canonical .txt (or .json) paths
above, annotated with the similarity score.
Search embedding
search needs a vector for the query, and the config decides where it comes
from. With embed set, a (text: string) => Promise<number[]> the caller
brings, the query is vectorized in-process, so a self-hosted Qdrant works and
mirage depends on no model runtime; examples/typescript/qdrant/qdrant_vfs.ts
feeds it a dependency-free hashed bag of words, and any model that returns a
vector plugs in the same way. Without it the query text goes to the server, so
the cluster must have inference enabled (Qdrant Cloud) and store vectors from
the same embeddingModel (default sentence-transformers/all-MiniLM-L6-v2).
The hook never lands in a snapshot: a restored mount asks for a fresh
VFS. Browsing (ls/cat/find/grep) needs no embedding.
Supported commands
ls, cd, tree, cat, stat, find, wc, head, tail, and search. grep/rg stay
lexical; search "<query>" <path> is the semantic command, returning ranked
points as canonical <id>.txt (or <id>.json) paths plus a score, so results
compose with cat, wc, and pipes. Flags: --top-k, --threshold, --method semantic.
Folder listings filter on payload fields. A filtered listing scrolls first and
only creates keyword payload indexes for the groupBy fields if Qdrant reports
one is required. maxRows caps how many points are scanned per folder.