Runtime inference has no cloud fallback and no network dependency.
Offline-first model optimization
AI that stays on your device.
Prepare Phi-3 Mini and Llama 3 artifacts for private, mobile-friendly local inference. No inference API key required. No hosted model API. No telemetry. No prompt collection.
Reproducible GGUF Q4/Q3 with checksum manifests; Phi-3 Q4_K_M verified. AWQ is optional and unverified.
Flutter FFI plugin with real token streaming. Android APK build verified; on-device phone and iPhone inference not yet verified.
Deployment ready. This site is a project overview; inference runs locally in the SDK, not in Vercel.