Llama.cpp on the Internet Computer

icpp · July 21, 2024, 4:30pm

This thread discusses llama.cpp on the Internet Computer.

A project funded by the DFINITY Grant: ICGPT V2

The first functioning version is now MIT licensed open source:

icpp · July 21, 2024, 4:31pm

Current status:

on a Mac, you can build/deploy/upload the LLM, and then call an endpoint to have it generate tokens.
it only works for a small model, because no use is made yet of the recent advancements of the IC (SIMD, float handling, etc.).
the README of the GitHub repo contains a list of TODOs

icpp · October 15, 2024, 6:21pm

It’s been a journey, but a pre-release of ICGPT with a llama.cpp backend is now live on the IC.

The deployed llama.cpp model is Qwen 2.5 - 0.5B - Q8_0 - Instruct
You can watch a small video here
You can try it out at https://icgpt.icpp.world/
A 0.5B model with q8_0 quantization fits fine in a 32bit canister.
However, because of the instruction limit, which requires multiple update calls, it takes about 2 minutes to get this answer to the question shown below
We did not do any load testing, so it will be interesting to see how it holds up when multiple users try it out at the same time.
The UI is still primitive. The same as the one that was developed for the tiny story teller LLM. Improving that is on the to-do list.

icpp · October 30, 2024, 4:47am

ICGPT V2 - final milestone reached

(I also posted this on X)

The grant work is now completed and here is a video summarizing what I created. I want to thank @dfinity for the opportunity & the support.

I also want to thank the #ICP community for the enthusiasm as I shared progress along the way over the past months, and the testing some of you did with the early releases.

Some of you even donated cycles that will keep the Qwen2.5 canister up & running for several months. You are the best.

You can try it out at: https://icgpt.icpp.world

I am very happy with the outcome of this project and there are big plans to build on top on this foundation. More on this later. But first some time to celebrate this milestone

Youtube Video

paulous · October 30, 2024, 5:49am

Worked better this time

superduper · November 7, 2024, 11:52pm

congrats super cool i hope this gets the use and attention it deserves

btw i’m playing with a solana project https://github.com/ai16z/eliza/ which can connect to to remote or local LLMs

does this have API end points so that we could add it as one of the remote models? i can only pidgeon code so hard for me to eval how it works

github.com

ai16z/eliza/blob/main/packages/core/src/core/models.ts

import settings from "./settings.ts";
import { Models, ModelProvider, ModelClass } from "./types.ts";

const models: Models = {
    [ModelProvider.OPENAI]: {
        endpoint: "https://api.openai.com/v1",
        settings: {
            stop: [],
            maxInputTokens: 128000,
            maxOutputTokens: 8192,
            frequency_penalty: 0.0,
            presence_penalty: 0.0,
            temperature: 0.6,
        },
        model: {
            [ModelClass.SMALL]: "gpt-4o-mini",
            [ModelClass.MEDIUM]: "gpt-4o",
            [ModelClass.LARGE]: "gpt-4o",
            [ModelClass.EMBEDDING]: "text-embedding-3-small",
        },

This file has been truncated. show original

icpp · November 8, 2024, 12:25pm

@superduper ,
thanks for your feedback and for pointing out the eliza model. I will check it out.

About the API:

There are two endpoints new_chat & run_update. These are candid based canister endpoints. The links bring you to the candid service definition. In the README it is described how you call it using dfx. This API is using the llama.cpp command line interface, and makes it really easy to test things locally, and then use the same arguments when calling the canister.
We’re looking at creating another API that is following the openAI completions standard.

roger · November 22, 2024, 4:50pm

Congratulations on the final milestones! Must have been some journey.

We just got done with our first milestone of LLM Marketplace and will be getting into an interesting phase. In this phase I was planning to explore Tiny llama (1.1 B parameters) and train it for a niche task.

But looking at your post I’m having second thoughts, primarily because your model is smaller than tiny llama and quantised. Despite, it hits the instruction limit restriction. I will experiment and share my learning.

I’m trying to brainstorm ideas to see what could be a tiny task it could be trained on bypassing the instruction limit.

Would appreciate your thoughts around it.

Cheers!

icpp · November 22, 2024, 8:38pm

Hi @roger ,

The 1.1B Tiny Llama will not fit.

I recommend you select a 0.5B parameter model, like the Qwen 2.5 model I am using, and try to fine tune that one.

roger · November 23, 2024, 11:21am

Yes, I have been contemplating to use Qwen and also exploring few other light wt models like SmolLM, DistilGpt2.

While researching these in also trying to close on a fun use case.

Thank you for the suggestion.

icpp · February 2, 2025, 3:16pm

We completed the update to the latest llama.cpp version (sha 615212).

Please start fresh by following the instructions at GitHub - onicai/llama_cpp_canister: llama.cpp for the Internet Computer

This update allows you run many new LLM architectures, including the 1.5Billion parameter DeepSeek model that attracted a lot of attention with this X post.

The main limiting factor of running the larger LLMs is the instructions limit. If a model can generate at least 1 token, you can use it, because we generate tokens via multiple update calls. (See the README in the repo for details.)

Latency is off course high, which hopefully will improve with further ICP protocol and hardware updates, but we believe it is already possible to build useful, targeted AI agents with their LLM running on-chain. It requires some smart prompt engineering, and this is an area where we are focusing our efforts.

To assist with prompt engineering, a python notebook prompt-design.ipynb is included in the repository, where you can run against the original llama.cpp compiled for your native system.

icpp · February 2, 2025, 3:38pm

A few notes on the testing we did with DeepSeek.

We tested this DeepSeek model, available on HuggingFace:

unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF · Hugging Face
the Q2_K model: DeepSeek-R1-Distill-Qwen-1.5B-Q2_K.gguf

The model card on Huggingface shows this llama.cpp command:

./llama.cpp/llama-cli \
    --model unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF/DeepSeek-R1-Distill-Qwen-1.5B-Q4_K_M.gguf \
    --cache-type-k q8_0 \
    --threads 16 \
    --prompt '<｜User｜>What is 1+1?<｜Assistant｜>' \
    -no-cnv

Our initial tests confirmed that the parameter --cache-type-k q8_0 is important to get a good answer from the Q2_K quantized model.

To call the canister, you would use something like this:

dfx canister call llama_cpp run_update '(record { args = vec {"--cache-type-k"; "q8_0"; "--prompt-cache"; "prompt.cache"; "--prompt-cache-all"; "-sp"; "-p"; "<｜User｜>What is 1+1?<｜Assistant｜>"; "-n"; "512" } })'

(No need to pass -no-cnv, because that is a default for llama_cpp_canister)

You can generate 2 tokens per update call, so you configure the LLM with this call (Details are in README of repo):

dfx canister call llama_cpp set_max_tokens '(record { max_tokens_query = 2 : nat64; max_tokens_update = 2 : nat64 })'

And that’s really it to get going with this model.

Topic		Replies	Views
Llama2.c LLM running in a canister! Programs & Applications	61	4864	July 1, 2024
Is llama-13B(or 7B) LLM possible to deploy on canister? Developers Discussing	4	499	June 6, 2024
How can I take an open source pretrained LLM model, deploy it to ICP and use as a private ChatGPT just fo me Developers	13	341	June 16, 2025
Introducing the LLM Canister: Deploy AI agents with a few lines of code Developers rust , DeAI	64	3878	July 30, 2025
AI and machine learning on the IC? Developers	114	10088	June 20, 2024

Llama.cpp on the Internet Computer

Related topics