---
name: 'Stepfun: Step 3.7 Flash (Free)'
slug: stepfun-step-3-7-flash
type: Multimodal
providers:
  - kilo
description: >-
  Step 3.7 Flash is StepFun's latest high-efficiency multimodal
  Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a
  vision encoder for native image and video understanding, activating roughly
  11B parameters per token. The model supports a 256K context window and exposes
  selectable reasoning levels (high/medium/low), letting callers trade off
  speed, cost, and depth of reasoning. Designed for coding, agentic workflows,
  structured outputs, and long-context productivity tasks.
tags:
  - multimodal
is_open_source: false
context_window: 262144
max_output: 262144
pricing:
  - image: 0
    input: 0
    output: 0
    provider: kilo (Free Tier)
release_date: '2024-01-01'
capabilities:
  audio: false
  vision: true
  tool_use: false
  citations: false
  reasoning: false
---
Stepfun: Step 3.7 Flash (Free) is a Multimodal model.
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token. The model supports a 256K context window and exposes selectable reasoning levels (high/medium/low), letting callers trade off speed, cost, and depth of reasoning. Designed for coding, agentic workflows, structured outputs, and long-context productivity tasks.
