IO

K2-Horizon-375B-A23B

LLMby Institute of Foundation Models·Model page

Institute of Foundation Models' open-weights 375B-A23B MoE flagship LLM from the K2-Horizon family.

Share:

Model Description

K2-Horizon-375B-A23B is the flagship of the K2-Horizon family: a sparse Mixture-of-Experts model that stores 375B parameters and runs 23B per token, with a 512K context window. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.

K2-Horizon-375B-A23B Highlights

  • Frontier-class agentic performance. On agentic tool use, terminal, and long-horizon workflow benchmarks it matches or beats open-weight MoE models up to 2.6× its size and is competitive with closed frontier models (see Benchmark Results).
  • 512K context. Native 524,288-token context from the midtraining stages onward.
  • Intermediate checkpoints. Intermediate checkpoints will be released so capability changes can be studied across training rather than at a single checkpoint.
  • Fully open. Training data/recipe and the training code will be made public.

Benchmark Results

1,4411,1621,2341,3801,4981,5691,5031,58434.014.229.115.334.631.128.737.365.334.345.553.759.967.564.871.625.38.012.820.526.233.528.034.724.89.019.023.826.928.625.431.767.745.751.248.872.466.974.065.372.844.477.183.5--83.3--84.750.934.252.356.455.050.460.0--Coding70.253.955.165.277.980.975.780.542.739.946.145.450.552.550.153.648.4--25.542.346.4------42.638.743.143.846.748.8----Scientific Reasoning32.028.431.939.041.139.538.541.387.386.787.292.989.591.189.691.18.63.15.43.720.921.022.916.9General76.071.073.380.376.778.373.377.023.023.042.017.024.043.045.040.074.770.032.082.074.07.010.061.0

Quickstart

Serving

vLLM, recipe at recipes.vllm.ai/IFM:

vllm serve IFM/K2-Horizon-375B-A23B \
  --revision main \
  --tensor-parallel-size 8 \
  --enable-expert-parallel \
  --trust-remote-code \
  --dtype bfloat16 \
  --max-model-len 131072 \
  --reasoning-parser k2_horizon \
  --tool-call-parser k2_horizon \
  --enable-auto-tool-choice

SGLang recipe validated on 8× H200 in the SGLang K2 Horizon cookbook:

python3 -m sglang.launch_server \
  --model-path IFM/K2-Horizon-375B-A23B \
  --revision main \
  --tp 8 \
  --ep 8 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --model-loader-extra-config '{"enable_multithread_load":false}' \
  --reasoning-parser k2_horizon \
  --tool-call-parser k2_horizon \
  --host 0.0.0.0 --port 30000

API Usage

[!Tip] Recommended settings: reasoning_effort="high", temperature=1.0, top_p=0.95. Reasoning depth is selected per request through chat_template_kwargs. Thinking is returned in reasoning_content and the answer in content.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="IFM/K2-Horizon-375B-A23B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)

Transformers

Validated with Transformers 4.57.6, PyTorch 2.13.0, Safetensors 0.8.0.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "IFM/K2-Horizon-375B-A23B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)

inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Best Practices

  1. Reasoning effort: always high. All reported results use high reasoning effort. Pass {"chat_template_kwargs": {"reasoning_effort": "high"}} on every request.
  2. Sampling parameters. temperature=1.0, top_p=0.95.
  3. Serving. Use the validated SGLang recipe above: BF16, TP=8 on one 8× H200 node, FlashAttention-3, wit
Author
IO
Institute of Foundation Models
Organization
IFM
Details
Downloads1.4K
Likes62
AccessOpen Source
Tasktext-generation
Parameters379B
Trending60
Licenseapache-2.0
Librarytransformers
CreatedSep 1, 2026
UpdatedSep 4, 2026
View on Hugging Face
Languages
en
Get the full context.

Sign up to read complete case studies, access detailed metrics, and unlock all use cases.

K2-Horizon-375B-A23B — AI Model Details | Applied