Quiz: On-device embedding models and runtimes

Why are small embedding models preferred for on-device RAG?