llamafu vs llama.rn: On-Device LLMs for Flutter and React Native

A practical comparison of llamafu, llama.rn and llama_cpp_dart for running llama.cpp models in mobile apps: what each binding covers and how to choose.

This comparison was rewritten in October 2026: an earlier version described llamafu features and published benchmarks that do not exist, and named a Flutter package we could not find.

The question

Should I use llamafu, llama.rn, or another llama.cpp binding to run an LLM in my mobile app?

The first answer is usually decided before any feature comparison: which framework is the app written in? llama.rn is a React Native binding. llamafu and llama_cpp_dart are Flutter plugins. If your app is React Native, the Flutter options are not options, and vice versa.

The 60-second version: all three wrap the same inference engine, llama.cpp, and all three are MIT-licensed. The differences are in the framework they target, the shape of the API they expose, and which llama.cpp capabilities they surface. None of them makes the underlying model faster than llama.cpp itself can run it.

What each project is

llamafu is a Flutter FFI plugin from Skelf Research. A native C++ layer wraps llama.cpp, Dart FFI bindings expose it, and a high-level async Dart API sits on top. It targets Android (API 21 and up) and iOS (12.0 and up), loads GGUF models from a file path, and exposes text generation with streaming, chat with conversation history, embeddings, vision models (LLaVA, Qwen2-VL), tool calling, schema-constrained JSON, GBNF grammar-constrained generation, and LoRA adapter loading and hot-swapping.

llama.rn is the React Native binding of llama.cpp, maintained at mybigday/llama.rn. Its README lists Metal acceleration on iOS and experimental Hexagon NPU support on Android, multimodal models through mmproj projectors, parallel decoding with slot-based request handling, tool calling through Jinja templates, GBNF and JSON-schema constrained output, and experimental on-device text-to-speech. From v0.10 it requires React Native’s New Architecture.

llama_cpp_dart is a community Dart FFI binding for llama.cpp, maintained at netdur/llama_cpp_dart. Its 0.9.x line is a rewrite focused squarely on llama.cpp inside a Flutter mobile app (iOS and Android, with macOS as a development target). Its README lists streaming, off-thread inference in a worker isolate, multimodal input, KV-cache persistence to a file, opt-in context shifting, speculative decoding, and Metal, CPU, OpenCL and Hexagon NPU builds. The author notes that the public API will likely have one more breaking change before 1.0.

The dimensions that matter

Dimensionllamafullama.rnllama_cpp_dart
App frameworkFlutterReact NativeFlutter
Enginellama.cppllama.cppllama.cpp
Model formatGGUFGGUFGGUF
PlatformsAndroid, iOSAndroid, iOSAndroid, iOS (macOS for development)
StreamingYesYesYes
Vision / multimodalYesYesYes
Tool callingYesYesNot a listed feature
Grammar / JSON-schema outputYesYesNot a listed feature
EmbeddingsYesYesSee project docs
LoRA adaptersYes, with hot-swapSee project docsSee project docs
KV-cache persistenceNot a listed featureSee project docsYes
Speculative decodingNot a listed featureYes (MTP, for models with MTP layers)Yes
LicenceMITMITMIT

“See project docs” means we have not verified the capability either way from the project’s README; check before relying on it. All three projects move quickly, so treat any table like this as a snapshot rather than a specification.

When to use which

Use llama.rn when:

  • Your app is React Native. That settles it.
  • You want parallel requests handled by the binding, or the experimental text-to-speech support.

Use llamafu when:

  • Your app is Flutter and you want a high-level, async Dart API where the common structured-output tasks — tool calls, schema-constrained JSON, grammar-constrained generation — are single method calls.
  • You want LoRA adapters you can load, apply and remove at runtime.
  • You are building an on-device agent and want the inference layer to already speak in tool calls and structured output.

Use llama_cpp_dart when:

  • Your app is Flutter and you need capabilities its README lists that llamafu’s does not, such as KV-cache persistence, context shifting or speculative decoding.
  • You want prebuilt native artefacts that include Hexagon NPU and OpenCL builds for Snapdragon devices.
  • You are comfortable tracking a binding that is still settling its public API.

What llamafu does NOT do

  • It does not publish benchmarks. We have not released per-device throughput, memory or quantisation-quality figures for llamafu. If you need numbers to defend a model-and-device choice, you will have to measure them on your target hardware — and record the device, OS version, llama.cpp revision, thread count and context size beside every result.
  • It is not desktop-first. It targets Android and iOS.
  • It does not ship a model hub. You bring your own GGUF file and pass its path. That is intentional — we believe in local models the developer controls — but it is worth knowing.
  • It does not change the physics. Memory is the binding constraint on mobile, and a binding cannot make a 7B model fit on a phone that does not have the RAM for it.

A realistic 30-minute evaluation

Fetch a small quantised model:

curl -L -o qwen2.5-1.5b-instruct-q4_k_m.gguf \
  https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct-GGUF/resolve/main/qwen2.5-1.5b-instruct-q4_k_m.gguf

Add the plugin with flutter pub add llamafu, put the model file somewhere the app can read it from the filesystem, and call it:

import 'package:llamafu/llamafu.dart';

Future<void> evaluate(String modelPath) async {
  final llamafu = await Llamafu.init(
    modelPath: modelPath,
    threads: 4,
    contextSize: 2048,
  );

  final text = await llamafu.complete(
    prompt: 'Say hello in three words.',
    maxTokens: 32,
    temperature: 0.7,
  );
  print(text);

  // Structured output is a single call
  final person = await llamafu.generateJson(
    prompt: 'Extract: John is 25 years old',
    schema: {
      'type': 'object',
      'properties': {
        'name': {'type': 'string'},
        'age': {'type': 'integer'},
      },
      'required': ['name', 'age'],
    },
  );
  print(person);

  llamafu.close();
}

Then do the part no comparison article can do for you: run the same model through the binding you are comparing, on the same device, with the same context size and thread count, and time both. Measure prompt processing and generation separately, and watch peak memory — the number that decides whether the OS kills your app.