Skip to content

[Workers AI] Update model reasoning types - #7437

Merged
danlapid merged 15 commits into
cloudflare:mainfrom
KastanDay:kastan/workers-ai-reasoning-types
Sep 25, 2026
Merged

danlapid merged 15 commits into
cloudflare:mainfrom
KastanDay:kastan/workers-ai-reasoning-types

Conversation

@KastanDay

@KastanDay KastanDay commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Refreshes the Workers AI binding types so each reasoning model's types accept exactly the reasoning controls it supports, with per-model hover docs.

This is an intentional TypeScript breaking change for code that passes a reasoning control the model does not support. Those requests were already ignored, remapped or rejected at runtime; now they fail to compile instead.

What changes

  • Exact efforts. reasoning_effort (and Responses reasoning.effort) accepts exactly the model's supported efforts from Workers AI ConfigAPI. Compatibility aliases (for example GLM-5.2 low, which maps to high) are not typed. The hover lists them as accepted by the API but not by the types.
  • No effort levels, no effort property. Gemma 4, GLM-4.7-Flash and Kimi K2.7 Code have no reasoning_effort. Their enable_thinking hover explains that.
Model reasoning_effort enable_thinking
GLM-5.2 max high none boolean
DeepSeek V4 Flash / Pro max high low none boolean
GLM-5.3, GLM-5.3-Flash max high low true
Kimi K2.6 high none boolean
Qwen3.8 27B low medium xhigh boolean
GPT-OSS 20B / 120B (Chat and Responses) low medium high true
Gemma 4, GLM-4.7-Flash none boolean
Kimi K2.7 Code none true
Nemotron 3 none (own chat_template_kwargs) boolean
Models without reasoning metadata shared low medium high shared boolean

Compatibility

Checked against main's generated-snapshot/index.ts, imported side by side:

  • No export is removed. New exports: the GLM-5.3, Nemotron Speech Streaming and renamed Gemma classes, and the Gemma, Nemotron 3 and Nemotron Speech Streaming input/output types.
  • postProcessedOutputs is identical for all 97 models in both.
  • inputs narrows for GPT-OSS 20B/120B, Kimi K2.6, Kimi K2.7 Code, GLM-5.2, DeepSeek V4 Flash/Pro, Qwen3.8 and GLM-5.3-Flash. Gemma 4, GLM-4.7-Flash and Nemotron 3 drop reasoning_effort, so object literals that pass it no longer compile. Every other model's inputs are unchanged.
  • The only AiModels key removed is img2img.

types/test/types/ai-reasoning.ts pins each model's exact effort union and covers supported calls, Nemotron's controls, and rejected values: unsupported efforts, aliases, arbitrary strings, the shared effort type on GLM-5.3, effort on models without levels, enable_thinking: false on mandatory models, and Responses minimal.

Generation

The declarations come from the SDK generator in cloudflare/ai/sdk!1778, run against production ConfigAPI. Only the reasoning, Nemotron and new-model declarations are applied to types/defines/ai.d.ts, because the rest of that file has drifted from the generator. generated-snapshot/ is the output of bazel build //types, and bazel test //types/... passes (29/29). The branch includes a merge of upstream/main.

The docs counterpart is cloudflare-docs#33541.

@KastanDay
KastanDay requested review from a team as code owners September 19, 2026 00:37
@KastanDay
KastanDay requested a review from cjol September 19, 2026 00:37
@github-actions

github-actions Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@KastanDay

KastanDay commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

I have read the CLA Document and I hereby sign the CLA

@KastanDay

Copy link
Copy Markdown
Contributor Author

recheck

github-actions Bot added a commit that referenced this pull request Sep 19, 2026
@KastanDay KastanDay changed the title Fix Workers AI model reasoning input types [Workers AI] Update model reasoning types Sep 21, 2026

@cjol cjol left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is definitely an improvement so I'm happy to merge, but it would be good to signpost to an authoritative "source of truth" (e.g. in a comment) to make it easy to keep these type definitions up-to-date in future. Where would that be?

@KastanDay

KastanDay commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

This is definitely an improvement so I'm happy to merge, but it would be good to signpost to an authoritative "source of truth" (e.g. in a comment) to make it easy to keep these type definitions up-to-date in future. Where would that be?

Thanks @cjol, I'll add a code comment with the source of truth: Workers AI ConfigAPI. That feeds a generator for both workerd types and dev docs.

One question that I've been asking internally is: "is it considered a breaking change to remove type hints?"

Today, ~all Workers AI models have the same type hints. In the future, we can show what they truly support.

Today: low, medium, high (same for all models). After: none, high, max.

The code still functions the same, i.e. we accept values outside the hints (compatibility aliases, below).

The alternative would be to show the fully compatibility list, but then we're back to showing the same thing for every model which isn't a very useful hint, e.g. none, minimal, low, medium, high, xhigh, max would be a bit silly.

Any opinion here?

…adata

Regenerated with the updated Workers AI type generator against production
model metadata:

- GLM 4.7 Flash only toggles reasoning, so it accepts no effort levels.
- Kimi K2.7 Code keeps the default efforts; reasoning cannot be disabled.
- Gemma 4 keeps its skip_special_tokens option alongside its efforts.
- Nemotron 3 uses the shared Chat Completions input that production publishes.
- Qwen 3.8 effort order follows its metadata.
- Name the default effort union ChatCompletionsReasoningEffort.
…strictions

Regenerated from the SDK generator (cloudflare/ai/sdk!1778) against
production model metadata. No code that type-checked before is rejected:

- `AiReasoningEffortHint<Suggested>` suggests a model's supported efforts
  in editors but accepts any string. Shared Chat Completions
  `reasoning_effort` and Responses request `reasoning.effort` use it too.
- Response types (`Reasoning`, `ReasoningEffort`, outputs) are unchanged.
- Models with reasoning metadata get `<Model>_ChatInput` /
  `_ChatTemplateKwargs` (and `_ResponsesInput` / `_Reasoning` for gpt-oss)
  interfaces extending the shared types with per-model docs: supported
  levels, aliases, defaults, and whether reasoning can be disabled.
- `enable_thinking` stays `boolean` for every model.
- Adds the public `@cf/zai-org/glm-5.3` model.
- Keeps `Base_Ai_Cf_Google_Gemma_4_26B_A4B_IT` as a deprecated alias of the
  generated `..._It` class.
Its metadata lists no supported efforts, so it has no effort levels even
though reasoning is always on. reasoning_effort still accepts any string.
Replace the exported per-model reasoning interfaces with unnamed types on
each model's Base_ class. Editor suggestions and per-model docs are
unchanged, and no new per-model type names are published, so editing
model metadata can never remove a name that user code references.
Regenerated from the Workers AI SDK type generator (cloudflare/ai/sdk!1778)
and production ConfigAPI. Reasoning types now enforce what each model
supports instead of suggesting values while accepting any string:

- reasoning_effort (and Responses reasoning.effort) accepts exactly the
  model's supported efforts; compatibility aliases are listed in the hover
  docs as accepted by the API but not by the types.
- Models without effort levels (Gemma 4, GLM-4.7-Flash, Kimi K2.7 Code) have
  no reasoning_effort; mandatory reasoning (GPT-OSS, Kimi K2.7 Code, GLM-5.3,
  GLM-5.3-Flash) types enable_thinking as true.
- Nemotron 3 publishes its own controls: no reasoning_effort, and
  chat_template_kwargs with enable_thinking, low_effort and
  force_nonempty_content.
- The shared ChatCompletionsInput and ResponsesInput keep main's types and
  still apply to models without reasoning metadata.
- Adds GLM-5.3 and Nemotron Speech Streaming; removes
  stable-diffusion-v1-5-img2img, which is no longer offered.

This intentionally rejects values these models do not support that the
types previously accepted.
@KastanDay

Copy link
Copy Markdown
Contributor Author

This is definitely an improvement so I'm happy to merge, but it would be good to signpost to an authoritative "source of truth" (e.g. in a comment) to make it easy to keep these type definitions up-to-date in future. Where would that be?

Thanks @cjol, I'll add a code comment with the source of truth: Workers AI ConfigAPI. That feeds a generator for both workerd types and dev docs.

One question that I've been asking internally is: "is it considered a breaking change to remove type hints?"

Today, ~all Workers AI models have the same type hints. In the future, we can show what they truly support.

Today: low, medium, high (same for all models). After: none, high, max.

The code still functions the same, i.e. we accept values outside the hints (compatibility aliases, below).

The alternative would be to show the fully compatibility list, but then we're back to showing the same thing for every model which isn't a very useful hint, e.g. none, minimal, low, medium, high, xhigh, max would be a bit silly.

Any opinion here?

The Workers AI team has decided to strictly enforce the correct Reasoning Effort types for our models. It's good to be correct. This PR only touches reasoning efforts, not other properties of the type system.

Instead of all models strictly enforcing "low | medium | high" reasoning, even when sometimes "medium" wasn't supported and a supported value of "max" was a type error. Now we strictly enforce each model's reasoning separately.

This PR is ready for review and merge! Thanks @cjol

@thatsKevinJain

Copy link
Copy Markdown
Contributor

The PR looks good to me ✅

@danlapid
danlapid merged commit ebd585e into cloudflare:main Sep 25, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants