Step 3.7 Flash
StepFun's multimodal reasoning model with a 256K-token context for text, image, and video inputs.
Design Arena
Design Arena ranks models on real-world front-end and design tasks — websites, UI components, data viz, SVG and more — through head-to-head human votes, scored as an ELO rating.| Category | ELO | Win rate | Rank |
|---|---|---|---|
| Asciiart | 1217 | 50.4% | #16 |
| Website | 1212 | 46.2% | #48 |
| UI component | 1206 | 43.6% | #49 |
| Code | 1204 | 44.8% | #51 |
| Game dev | 1200 | 41.2% | #48 |
Pricing
per 1M tokensInput$0.2 /1M
Output$1.15 /1M
Cache read$0.04 /1M
Model Description
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Get the full context.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.
Get the full context.
Sign up to read complete case studies, access detailed metrics, and unlock all use cases.