Stepfun: Step 3.7 Flash (Free)
Stepfun: Step 3.7 Flash (Free) is a Multimodal model.
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token. The model supports a 256K context window and exposes selectable reasoning levels (high/medium/low), letting callers trade off speed, cost, and depth of reasoning. Designed for coding, agentic workflows, structured outputs, and long-context productivity tasks.
Vision
Audio
Tool Use
Reasoning
Citations
Key Specs
Context
262k
Max Output
262k
Best Price (Input / Output)
Chat with Model
Free / Free per 1M tokens
Pricing Comparison
| Provider | Input (1M) | Output (1M) | Image (1k) |
|---|---|---|---|
| kilo (Free Tier) | Free | Free | - |
Prices are per 1 million tokens unless otherwise noted. Image pricing is per 1000 images if applicable.
Price History
Prompt Price
Completion Price
Price trend per 1 million tokens over time.