S

Step 3.7 Flash

Multimodalby StepFun·Model page

StepFun's multimodal reasoning model with a 256K-token context for text, image, and video inputs.

Rankings
#11on OpenRouter39.6%
5395.2B tokens
View rankings
Max output

Most tokens the model can return in a single response.

256Ktokens
Reasoning
Reasoning-first model

Works through a step-by-step chain of thought before answering.

Share:

Design Arena

Design Arena ranks models on real-world front-end and design tasks — websites, UI components, data viz, SVG and more — through head-to-head human votes, scored as an ELO rating.
CategoryELOWin rateRank
Asciiart121750.4%#16
Website121246.2%#48
UI component120643.6%#49
Code120444.8%#51
Game dev120041.2%#48

Pricing

per 1M tokens
Input$0.2 /1M
Output$1.15 /1M
Cache read$0.04 /1M

Model Description

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Author
S
StepFun
Organization
stepfun
Details
Downloads
Likes
AccessOpen Source
Context256K tokens
Input price$0.2 /1M
Output price$1.15 /1M
CreatedMay 28, 2026
Updated
View on Hugging Face
Benchmarks
Intelligence30.3
Coding39.6
Agentic21.5
Get the full context.

Sign up to read complete case studies, access detailed metrics, and unlock all use cases.

Step 3.7 Flash — AI Model Details | Applied