On Device AI
Ground local AI with instant access to user context.
Semantic, keyword, and hybrid search that runs entirely on device. Embed, index, and retrieve in under 10ms with no cloud infrastructure, no vector database, and no network round trip.
Your AI app
Retrieval Runtime powered by Moss
<10ms
Retrieval
32 MB
One Model
100%
Offline
CPU Only
No NPU Required
Used in production by companies backed by Y-Combinator & a16z
Problem
Modern AI runs on phones, laptops, browsers, vehicles, wearables, and edge devices. Retrieval still runs in the cloud, adding latency, infrastructure cost, and privacy tradeoffs.
Local AI needs a different retrieval architecture.
Moss brings semantic, keyword, and hybrid search directly to where your AI runs, eliminating cloud latency, reducing infrastructure costs, and keeping data on device.
Ground local AI with instant access to user context.
Search everything inside your application by meaning, not keywords.
Retrieve context across files, notes, photos, messages, and application data.
Runs entirely on CPU with no cloud infrastructure or external services.
Messages1ms
100K+ msgs
SMS, chat, attachment context
Files1ms
10K+ files
Document text + filenames + metadata
Gallery



Notes
offsite agenda — day 1, Fort Mason
Messages
…flights into SFO are booked for the offsite!
Files
Offsite plan June 20.docx
Offsite plan June 20 (2).pdf
Gallery3ms
50K+ photos
Image captions + EXIF + people tags
Notes0.8ms
100K+ notes
Notes body + handwriting OCR
Moss runs the entire retrieval pipeline on device.
Whether you're powering AI assistants, semantic search, RAG, recommendations, memory, or system search, it all happens locally.
No cloud APIs.
No embedding service.
No vector database.
No round trip.
Just instant retrieval where your users already are.
Moss engine
Your AI app
finish the investor update
I found last month's update, the MRR chart, and the Leads from yesterday.
I'll draft this month's version with the same structure and update the metrics.
Everything developers need in one SDK.
One lightweight model for every modality.
Purpose built for edge hardware.
Ground local AI in documents, messages, notes, photos, and app data. Nothing leaves the device.
Search every piece of user content inside your app. Built for productivity, media, ecommerce, travel, finance, and health.
One local retrieval engine across phones, vehicles, TVs, wearables, and embedded devices. System-wide search, fast and offline.
Secure AI search that keeps sensitive company knowledge on employee devices.
| System | Average Search Latency |
|---|---|
| Moss | 3.1 ms |
| ChromaDB | 351.8 ms |
| Pinecone | 432.6 ms |
| Qdrant | 597.6 ms |
Benchmarks include embedding inference and retrieval across 100,000 documents. Most competing benchmarks measure retrieval only.
Embed. Index. Retrieve. All on the device.