Spark-X2.5-4B-GGUF
SparkLLM's GGUF quantizations of Spark-X2.5-4B for local use with llama.cpp, Ollama and LM Studio.
Base model
Model Description
[!NOTE] This repository provides a BF16 GGUF conversion of Spark-X2.5-4B.
Spark-X2.5 is a compact, general-purpose language model for conversation, writing, translation, reasoning, coding, tool use, and agentic workflows. It uses a hybrid attention architecture, supports a native context length of up to 1M tokens, and covers more than 200 languages. For its architecture, training methods, benchmark results, fine-tuning, and citation, see the Spark-X2.5-4B.
Local Deployment
The GGUF file can be used for local inference with Ollama and LM Studio. Spark-X2.5 support is provided by XHToken/llama.cpp, so the Quick Starts below use this compatible implementation.
Ollama Quick Start
Build
git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
git clone https://github.com/ollama/ollama.git ollama-spark
cd ollama-spark
export OLLAMA_LLAMA_CPP_SOURCE="$(cd ../llama.cpp-spark && pwd)"
cmake -S . -B build
cmake --build build --parallel 8
Import the GGUF
Replace the model path below with the absolute path to the downloaded GGUF file:
printf 'FROM /absolute/path/to/Spark-X2.5-4B.gguf\n' > ./Modelfile.spark
Create and Run
Start the Ollama server in the first terminal:
./ollama serve
Open a second terminal in the same ollama-spark directory:
./ollama create Spark-X2.5-4B -f ./Modelfile.spark
./ollama run Spark-X2.5-4B --think=false
--think=false disables thinking mode for faster, direct responses.
LM Studio Quick Start
Build the Compatible llama.cpp Runtime
git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
cd llama.cpp-spark
cmake -S . -B build
cmake --build build --parallel 8
Configure LM Studio
Close LM Studio.
Back up the selected LM Studio runtime directory:
<LM_STUDIO_HOME>/extensions/backends/<selected-runtime>/Copy the
llama.cpp-sparkbuild output into the selected runtime directory, replacing the existing runtime files.Place
Spark-X2.5-4B.ggufin:<LM_STUDIO_HOME>/models/<org>/<name>/
Example runtime directory on Apple Silicon:
./build/bin/* -> ~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-<version>/
Run
Open LM Studio, select the model under My Models, click Load, and start a new chat.
You can also use the lms CLI:
lms ls
lms load <model>
lms chat <model>
License
Released under the Apache License 2.0.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.