The New Homepage: An AI Agent Live Session@Thu Oct 1 · 9am PT

Join us
Moss
usemossStart Free
Blog/Engineering

Why We Built a Search Runtime in Rust and Compiled It to WebAssembly

Date

September 25, 2026

Written by

  • Ashvath Suresh Kumar, Founding Growth
The word RUST drawn in gray pixel blocks above the word WASM in moss-green pixel blocks, separated by a dashed line.
  • How Moss Separates Index Building from Queries
  • Why Rust and WebAssembly Fit the Runtime
  • What Happens When an Index Loads and Changes
  • What the Published Performance Results Show
  • Where Portability Adds Engineering Work
  • Keeping Retrieval with the Agent

We built Moss around the idea that retrieval belongs in the runtime with the agent. When an agent needs knowledge to answer a question, it should be able to search that knowledge where it is already running, whether that is a browser tab, a server, a phone, or an edge environment. We wanted retrieval to be a capability the agent carries with it, without adding a network service to every lookup.

That decision came from the applications we were building for. A voice agent checking a refund policy has to retrieve the answer while a caller waits. A website assistant may need to search inside the browser, and a field application may need access to its knowledge after the connection drops. These applications have different deployment constraints, but they all benefit when retrieval can live beside the code that uses it.

Rust and WebAssembly followed from that requirement. Rust gave us control over execution and memory in an embedded library, while WebAssembly made that library available inside the browser. The more consequential design choice was deciding which work belonged in the cloud and which work had to happen in the agent’s process.

How Moss Separates Index Building from Queries

Moss Cloud handles the work of preparing and distributing an index. It ingests documents, builds the searchable knowledge, and makes updates available to the runtime. The application loads the index locally before it starts querying. Once it is loaded, the runtime can embed a query, retrieve relevant documents, apply filtering and ranking, and return context to the agent without a retrieval request to the cloud. The offline search documentation describes this lifecycle.

This split keeps index maintenance available as a managed service while making query execution part of the application. It also changes the failure boundary. After the required assets and index are loaded, a temporary loss of connectivity does not require every knowledge lookup to fail, although receiving new content still requires synchronization.

The diagram shows both parts of the architecture. The build targets determine how the Rust core runs on each platform. Index delivery is a separate path from Moss Cloud to the deployed runtime, where the agent and its loaded knowledge meet.

Two architectures compared. Traditional retrieval: the agent makes a network request to a vector database for each query. Moss architecture: queries run inside the agent's runtime. A shared Rust core is compiled as a WebAssembly build, deployed to the browser, and a native build, deployed to server, mobile, and edge. In each, the agent talks to a local runtime with a loaded index. Moss Cloud builds and syncs indexes, and a dashed path delivers the index and background refreshes to both runtimes.

The solid arrows show compilation and deployment. The dashed path delivers index updates outside the per-query retrieval path.

Why Rust and WebAssembly Fit the Runtime

An embedded search engine shares a process with the application using it. Its memory usage and execution time therefore become part of that application’s behavior. Rust was a good fit because it lets us manage memory without adding a garbage collector to the search core, while its ownership system prevents many classes of memory errors in safe code. That gives us useful control over latency without making memory safety entirely a matter of programmer discipline. Rust’s ownership model explains the underlying tradeoff.

Keeping retrieval behavior in the Rust core also means a ranking change can reach the supported SDKs through the same implementation. Platform bindings handle how applications call the runtime and receive results. They do not need to recreate the search logic, which makes it easier to reason about correctness as the runtime moves between environments.

WebAssembly gives the browser a way to execute that core within its own security and execution model. Native builds serve hosts that can load a platform library directly. These targets can expose a similar application workflow, but they still have different startup costs and access to system resources.

Sharing the implementation helps keep behavior aligned; it does not imply identical performance on a laptop, a phone, and a browser tab.

What Happens When an Index Loads and Changes

A local query starts being useful only after the runtime is ready, so startup deserves its own budget. In the browser, first use includes downloading the WebAssembly module and embedding model, as described in the browser API reference. The application also has to fetch and load its index.

Network transfer, runtime initialization, and index loading all contribute to the time before the first query can run. A query latency measurement by itself does not describe that experience.

Loading an index does not require embedding the document collection again. That work has already happened during the cloud build. Caching can also reduce repeated transfer: the native JavaScript SDK can keep an index on disk and reuse it on a later load when the cloud version has not changed. This avoids fetching the same data again, though the application still needs a ready runtime and a loaded index. The storage documentation covers that distinction.

Index size matters in several ways. The downloaded bytes affect transfer time and local storage, while the loaded index consumes memory alongside the embedding model and the application itself. Text, metadata, and temporary query buffers also contribute to the process’s footprint. Document count alone is therefore an incomplete sizing guide: a collection of short support answers and a collection of long manuals can impose different costs even when they contain the same number of records.

Updates have a lifecycle too. With automatic refresh enabled, the runtime checks for a newer index and downloads a ready version in the background. Queries already in progress can finish against the current version while later calls use the replacement. That behavior, documented in data hydration and sync, keeps refresh work from interrupting an active query.

It also means memory planning should leave room for an update while the old version is still serving requests, rather than treating steady-state usage as the maximum.

What the Published Performance Results Show

Our repository reports an end-to-end query benchmark over 100,000 documents, with 750 measured queries returning the top five results on a MacBook Pro with an M4 Pro and 24 GB of memory. The measurement includes query embedding and search. The published results give the following latency distribution.

WorkloadMedianP95P99
Embedding and search3.1 ms4.3 ms5.4 ms

For that setup, even the 99th percentile remained below six milliseconds. The scope matters: this is a query benchmark, and it does not establish download time, index loading time, or browser performance. It shows that the local runtime can fit retrieval into a small part of the response budget on the tested machine.

A separate voice-agent pipeline profile reports the search processing step inside a hosted retrieval flow and a local one. The hosted database processing step had a median of 18 ms; the in-process search step had a median of 1.2 ms. Those figures exclude query embedding and result processing, so they describe a narrower part of the request than the repository benchmark.

The pipeline profile does not specify its hardware, corpus size, or query count, so it is an illustrative result for that workload rather than a reproducible comparison across platforms. Taken with the repository benchmark, it gives us evidence for local execution while leaving native-versus-WebAssembly performance, startup time, and peak memory to be measured on the application’s actual deployment target.

Where Portability Adds Engineering Work

The browser makes the cost of distribution immediately visible. A larger module or model increases the work before retrieval is ready, and loading an index consumes resources in a tab that also has to render the interface and run the rest of the application. A deployment decision therefore has to account for the first visit as well as repeated queries. An index that works comfortably beside a server agent may need a smaller scope when it is delivered to a browser on a constrained device.

Browser execution also depends on the page that hosts it. For example, shared memory between workers requires a secure context and cross-origin isolation under the browser’s shared-memory rules. That is a constraint on the capabilities a library can assume when embedded in an arbitrary site.

Serving WebAssembly brings its own integration details: streaming instantiation expects the appropriate content type, and a site’s content security policy can restrict compilation, as the WebAssembly documentation explains. A successful build alone does not establish that the surrounding page can run it correctly.

Native distribution moves the work elsewhere. A native library has to match the operating system, processor architecture, and host language that will load it. Linux compatibility is especially easy to miss when development and deployment environments differ. Keeping retrieval in one core reduces duplicated search code, but each supported package still needs to load and behave correctly in its host environment. Compatibility is part of the runtime’s usefulness, alongside the time it takes to answer a query.

Local execution also makes resource ownership explicit. The application pays for the CPU and memory used by retrieval, and the index it loads needs to fit the device’s budget. Where the corpus is too large to distribute to clients, running Moss beside a server-based agent can be a better placement. The architecture gives us a way to choose that location without changing the application into a client of a separate search service on every turn.

Keeping Retrieval with the Agent

The engineering work follows from the reason we built Moss this way. Retrieval belongs where the agent runs, with the knowledge it needs available when it needs to act.

Rust gives us the foundation for an embedded implementation, and WebAssembly carries it into the browser. Together with cloud index building and synchronization, they let us keep the work of answering a query close to the agent while managing the knowledge it searches over time.

You can try Moss or explore the documentation to see how the runtime fits into your application.

Read more

  • Why AI Infrastructure Is Moving Into the Runtime

    EngineeringSeptember 9

  • What Happens When You Remove the Network Hop from RAG

    EngineeringMarch 17

  • Building Voice AI That Feels Human: A Latency Budget Breakdown

    EngineeringAugust 19

Ready to ship faster AI products?

Moss gives you production-ready semantic retrieval without infrastructure complexity.

Test performanceTalk to an Engineer
Moss
AICPA SOC 2 Type 2HIPAA

Product

Founding AgentLocal Search

Use Cases

Voice AIAI CopilotsIn-App SearchOn-Device AI

Company

PricingBlogCareersBrand Kit

Resources

DocsGlossaryBenchmarks

Integrations

DSPyElevenLabsLangChainLiveKitMCP ServerNext.jsPipecatVAPIVercel AI SDKVitePress

© 2026 MOSS

Privacy PolicyTerms of ServiceTrust Center