Device: Vivo 1820 (2GB RAM)
OS: Android 8.1 (Oreo)
Model: Gemma3 270M IT (Quantized)
PROCESSOR: Helio P22 (64-bit)
Status: RUNNING Download the apk File
demo_video.mp4
Hi there! This isn't just a regular Flutter app. This represents a stubborn developer's refusal to accept "your device is too old" as an answer.
I wanted to run a Local Large Language Model (LLM) entirely offline. No internet, no APIs, just raw compute on my old Vivo 1820. I tried a lot of things. And oh boy, did it crash. A lot.
I started with the standard llama_cpp_dart package. It compiled, installed, and... SIGSEGV. Instant crash. The app would just vanish the moment I tried to load the model.
Most people would assume "out of memory" on a 2GB phone. But the logs showed something weirder. It was crashing before the model even loaded into RAM. It was crashing during the "handshake" between my Dart code and the C++ engine.
I dug into the native C++ code (llama.h) and compared it with the Dart FFI (Foreign Function Interface) bindings. It was like comparing two blueprints for the same house, but one was missing a room.
The Dart code was allocating a memory block for llama_context_params that was too small. It didn't know about the new samplers field that the C++ engine expected. So when the C++ engine tried to write its default settings effectively...
- It filled the struct.
- It kept writing past the end of the allocated memory.
- BOOM. Memory corruption.
I had to manually patch the library on my machine to fix this. Here's exactly what I did to make this potato fly:
- Patched the FFI Bindings: I physically edited
llama_cpp.dartto add the missingsamplersanduse_direct_iofields. This aligned the "blueprints" perfectly. - Forced 32-bit Architecture: My phone is old. Modern libraries default to 64-bit. I had to strictly filter the
ndkconfiguration toarmeabi-v7aandarm64-v8ato stop it from trying to load incompatible x86 libraries. - The "Android 8.1" Special: Older Androids handle shared libraries differently. I had to manually ensure
libomp.so(OpenMP) was loaded explicitly before the main library, or the linker would give up.
If you want to run this code (especially on an older device)like mine, you can't just flutter run. You need my specific patches.
Check out the offline_llm_integration_guide.md in this repo for the exact lines of code you need to change in the llama_cpp_dart package.
It's not ChatGPT-4o fast. It takes a few seconds to think. But it works. It's a neural network thinking inside a 7500 rupee phone from 2020.
And that, to me, is magic.
This project is running on Vivo 1820, proving that AI can run anywhere.
- By a curious 'me its me' who likes fixing things.*





