Skip to main content
CrashBytes Training
Home
/
on-device-mobile-rag
On-Device and Mobile RAG
Why run RAG on-device
On-device embedding models and runtimes
Quantization and model size
On-device vector store and ANN
First-launch indexing and freshness
Latency, memory, and battery budgets
Where does generation run?
Mobile RAG architecture pattern
Updating models and data on mobile