llamafu vs llama.rn: On-Device LLMs for Flutter and React Native
A practical comparison of llamafu, llama.rn and llama_cpp_dart for running llama.cpp models in mobile apps: what each binding covers and how to choose.
This comparison was rewritten in October 2026: an earlier version described llamafu features and published benchmarks that do not exist, and named a Flutter package we could not find.
The question
Should I use llamafu, llama.rn, or another llama.cpp binding to run an LLM in my mobile app?
The first answer is usually decided before any feature comparison: which framework is the app written in? llama.rn is a React Native binding. llamafu and llama_cpp_dart are Flutter plugins. If your app is React Native, the Flutter options are not options, and vice versa.
The 60-second version: all three wrap the same inference engine, llama.cpp, and all three are MIT-licensed. The differences are in the framework they target, the shape of the API they expose, and which llama.cpp capabilities they surface. None of them makes the underlying model faster than llama.cpp itself can run it.
What each project is
llamafu is a Flutter FFI plugin from Skelf Research. A native C++ layer wraps llama.cpp, Dart FFI bindings expose it, and a high-level async Dart API sits on top. It targets Android (API 21 and up) and iOS (12.0 and up), loads GGUF models from a file path, and exposes text generation with streaming, chat with conversation history, embeddings, vision models (LLaVA, Qwen2-VL), tool calling, schema-constrained JSON, GBNF grammar-constrained generation, and LoRA adapter loading and hot-swapping.
llama.rn is the React Native binding of llama.cpp, maintained
at mybigday/llama.rn. Its README lists Metal acceleration on iOS
and experimental Hexagon NPU support on Android, multimodal models
through mmproj projectors, parallel decoding with slot-based request
handling, tool calling through Jinja templates, GBNF and JSON-schema
constrained output, and experimental on-device text-to-speech. From
v0.10 it requires React Native’s New Architecture.
llama_cpp_dart is a community Dart FFI binding for llama.cpp,
maintained at netdur/llama_cpp_dart. Its 0.9.x line is a rewrite
focused squarely on llama.cpp inside a Flutter mobile app (iOS and
Android, with macOS as a development target). Its README lists
streaming, off-thread inference in a worker isolate, multimodal
input, KV-cache persistence to a file, opt-in context shifting,
speculative decoding, and Metal, CPU, OpenCL and Hexagon NPU
builds. The author notes that the public API will likely have one
more breaking change before 1.0.
The dimensions that matter
| Dimension | llamafu | llama.rn | llama_cpp_dart |
|---|---|---|---|
| App framework | Flutter | React Native | Flutter |
| Engine | llama.cpp | llama.cpp | llama.cpp |
| Model format | GGUF | GGUF | GGUF |
| Platforms | Android, iOS | Android, iOS | Android, iOS (macOS for development) |
| Streaming | Yes | Yes | Yes |
| Vision / multimodal | Yes | Yes | Yes |
| Tool calling | Yes | Yes | Not a listed feature |
| Grammar / JSON-schema output | Yes | Yes | Not a listed feature |
| Embeddings | Yes | Yes | See project docs |
| LoRA adapters | Yes, with hot-swap | See project docs | See project docs |
| KV-cache persistence | Not a listed feature | See project docs | Yes |
| Speculative decoding | Not a listed feature | Yes (MTP, for models with MTP layers) | Yes |
| Licence | MIT | MIT | MIT |
“See project docs” means we have not verified the capability either way from the project’s README; check before relying on it. All three projects move quickly, so treat any table like this as a snapshot rather than a specification.
When to use which
Use llama.rn when:
- Your app is React Native. That settles it.
- You want parallel requests handled by the binding, or the experimental text-to-speech support.
Use llamafu when:
- Your app is Flutter and you want a high-level, async Dart API where the common structured-output tasks — tool calls, schema-constrained JSON, grammar-constrained generation — are single method calls.
- You want LoRA adapters you can load, apply and remove at runtime.
- You are building an on-device agent and want the inference layer to already speak in tool calls and structured output.
Use llama_cpp_dart when:
- Your app is Flutter and you need capabilities its README lists that llamafu’s does not, such as KV-cache persistence, context shifting or speculative decoding.
- You want prebuilt native artefacts that include Hexagon NPU and OpenCL builds for Snapdragon devices.
- You are comfortable tracking a binding that is still settling its public API.
What llamafu does NOT do
- It does not publish benchmarks. We have not released per-device throughput, memory or quantisation-quality figures for llamafu. If you need numbers to defend a model-and-device choice, you will have to measure them on your target hardware — and record the device, OS version, llama.cpp revision, thread count and context size beside every result.
- It is not desktop-first. It targets Android and iOS.
- It does not ship a model hub. You bring your own GGUF file and pass its path. That is intentional — we believe in local models the developer controls — but it is worth knowing.
- It does not change the physics. Memory is the binding constraint on mobile, and a binding cannot make a 7B model fit on a phone that does not have the RAM for it.
A realistic 30-minute evaluation
Fetch a small quantised model:
curl -L -o qwen2.5-1.5b-instruct-q4_k_m.gguf \
https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct-GGUF/resolve/main/qwen2.5-1.5b-instruct-q4_k_m.gguf
Add the plugin with flutter pub add llamafu, put the model file
somewhere the app can read it from the filesystem, and call it:
import 'package:llamafu/llamafu.dart';
Future<void> evaluate(String modelPath) async {
final llamafu = await Llamafu.init(
modelPath: modelPath,
threads: 4,
contextSize: 2048,
);
final text = await llamafu.complete(
prompt: 'Say hello in three words.',
maxTokens: 32,
temperature: 0.7,
);
print(text);
// Structured output is a single call
final person = await llamafu.generateJson(
prompt: 'Extract: John is 25 years old',
schema: {
'type': 'object',
'properties': {
'name': {'type': 'string'},
'age': {'type': 'integer'},
},
'required': ['name', 'age'],
},
);
print(person);
llamafu.close();
}
Then do the part no comparison article can do for you: run the same model through the binding you are comparing, on the same device, with the same context size and thread count, and time both. Measure prompt processing and generation separately, and watch peak memory — the number that decides whether the OS kills your app.
What to read next
- Running Language Models on Your Phone: The llamafu Experiment — the memory arithmetic and what we would measure
- Autonomous Mobile Agents: ukkin’s Architecture for On-Device AI — the agent layer
- llamafu repository
- llama.rn repository
- llama_cpp_dart repository