MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF
GnLOLot's GGUF quantization of a 1B MiniCPM5 model fine-tuned for thinking and tool-calling.
Base model
Model Description
MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF
GGUF quantizations of MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.
This repository provides local-deployment builds of a 1B Thinking model fine-tuned on Fable 5 data (V2) atop openbmb/MiniCPM5-1B. Compared with V1, V2 strengthens tool calling / function calling, while keeping MiniCPM5's native chat template embedded in the GGUF files.
Transformers checkpoint: MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking
Previous GGUF version: MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF (V1)
Files
| File | Quant | Size | Notes |
|---|---|---|---|
MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.gguf |
Q8_0 | ~1.1 GB | recommended default |
MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-F16.gguf |
F16 | ~2.1 GB | full-precision conversion base |
Q8_0 is the recommended default quant for this 1B model.
Quick start
llama.cpp (llama-cli)
llama-cli \
-m MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.gguf \
-p "Write a Python function to merge two sorted lists." \
-n 512 \
--temp 0.9 --top-p 0.95 \
-c 8192
The model supports up to 128K tokens (131,072) per
config.json. Set-caccording to your available VRAM/RAM.
llama.cpp server
llama-server \
-m MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.gguf \
-c 8192 --port 8080
LM Studio / jan / KoboldCpp
Load any .gguf file from this repository. The MiniCPM5 chat template is embedded in the GGUF metadata.
Sampling recommendations
Generation defaults are inherited from MiniCPM5-1B:
| Mode | Params |
|---|---|
| Think (default) | temperature=0.9, top_p=0.95 |
| No Think | temperature=0.7, top_p=0.95, enable_thinking=False |
Capabilities
- Tool calling (enhanced in V2) — stronger function-calling / tool-use behavior
- Fable 5 fine-tune (V2) — post-trained on Fable 5 data
- Coding — code generation, debugging, and software-engineering workflows
- Instruction following — more reliable adherence to user prompts and task constraints
- Thinking mode — chain-of-thought reasoning; MiniCPM5 chat template baked into the GGUF
- Long context — up to 128K tokens (131,072 tokens per upstream
config.json)
Benchmark
Scores for the Transformers checkpoint MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking:
BFCL + API-Bank
| Model | BFCL non_live | BFCL live | API-Bank |
|---|---|---|---|
| MiniCPM5-1B (Base) | 41.51% | 60.24% | 7.30% |
| MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking | 43.06% | 63.33% | 22.10% |
Tau-Bench
| Domain | MiniCPM5-1B (Base) | MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking |
|---|---|---|
| Airline | 0.34 (17/50) | 0.36 (18/50) |
| Retail | 0.052 (6/115) | 0.070 (8/115) |
Limitations
- Thinking outputs — the model may emit reasoning blocks before the final answer
- 1B scale — lightweight local deployment; not frontier-scale
- Runtime context — actual usable context depends on your GGUF runtime and hardware limits
Provenance & licensing
Apache-2.0, inherited from MiniCPM5-1B.
Acknowledgements
- Base model: OpenBMB / MiniCPM5-1B
- Transformers checkpoint: MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking
- Quantization: llama.cpp
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.