[Workers AI] Update model reasoning types - #7437
Conversation
|
All contributors have signed the CLA ✍️ ✅ |
|
I have read the CLA Document and I hereby sign the CLA |
|
recheck |
cjol
left a comment
There was a problem hiding this comment.
This is definitely an improvement so I'm happy to merge, but it would be good to signpost to an authoritative "source of truth" (e.g. in a comment) to make it easy to keep these type definitions up-to-date in future. Where would that be?
Thanks @cjol, I'll add a code comment with the source of truth: Workers AI ConfigAPI. That feeds a generator for both workerd types and dev docs. One question that I've been asking internally is: "is it considered a breaking change to remove type hints?" Today, ~all Workers AI models have the same type hints. In the future, we can show what they truly support. Today: The code still functions the same, i.e. we accept values outside the hints (compatibility aliases, below). The alternative would be to show the fully compatibility list, but then we're back to showing the same thing for every model which isn't a very useful hint, e.g. Any opinion here? |
…adata Regenerated with the updated Workers AI type generator against production model metadata: - GLM 4.7 Flash only toggles reasoning, so it accepts no effort levels. - Kimi K2.7 Code keeps the default efforts; reasoning cannot be disabled. - Gemma 4 keeps its skip_special_tokens option alongside its efforts. - Nemotron 3 uses the shared Chat Completions input that production publishes. - Qwen 3.8 effort order follows its metadata. - Name the default effort union ChatCompletionsReasoningEffort.
…strictions Regenerated from the SDK generator (cloudflare/ai/sdk!1778) against production model metadata. No code that type-checked before is rejected: - `AiReasoningEffortHint<Suggested>` suggests a model's supported efforts in editors but accepts any string. Shared Chat Completions `reasoning_effort` and Responses request `reasoning.effort` use it too. - Response types (`Reasoning`, `ReasoningEffort`, outputs) are unchanged. - Models with reasoning metadata get `<Model>_ChatInput` / `_ChatTemplateKwargs` (and `_ResponsesInput` / `_Reasoning` for gpt-oss) interfaces extending the shared types with per-model docs: supported levels, aliases, defaults, and whether reasoning can be disabled. - `enable_thinking` stays `boolean` for every model. - Adds the public `@cf/zai-org/glm-5.3` model. - Keeps `Base_Ai_Cf_Google_Gemma_4_26B_A4B_IT` as a deprecated alias of the generated `..._It` class.
Its metadata lists no supported efforts, so it has no effort levels even though reasoning is always on. reasoning_effort still accepts any string.
Replace the exported per-model reasoning interfaces with unnamed types on each model's Base_ class. Editor suggestions and per-model docs are unchanged, and no new per-model type names are published, so editing model metadata can never remove a name that user code references.
Regenerated from the Workers AI SDK type generator (cloudflare/ai/sdk!1778) and production ConfigAPI. Reasoning types now enforce what each model supports instead of suggesting values while accepting any string: - reasoning_effort (and Responses reasoning.effort) accepts exactly the model's supported efforts; compatibility aliases are listed in the hover docs as accepted by the API but not by the types. - Models without effort levels (Gemma 4, GLM-4.7-Flash, Kimi K2.7 Code) have no reasoning_effort; mandatory reasoning (GPT-OSS, Kimi K2.7 Code, GLM-5.3, GLM-5.3-Flash) types enable_thinking as true. - Nemotron 3 publishes its own controls: no reasoning_effort, and chat_template_kwargs with enable_thinking, low_effort and force_nonempty_content. - The shared ChatCompletionsInput and ResponsesInput keep main's types and still apply to models without reasoning metadata. - Adds GLM-5.3 and Nemotron Speech Streaming; removes stable-diffusion-v1-5-img2img, which is no longer offered. This intentionally rejects values these models do not support that the types previously accepted.
The Workers AI team has decided to strictly enforce the correct Reasoning Effort types for our models. It's good to be correct. This PR only touches reasoning efforts, not other properties of the type system. Instead of all models strictly enforcing "low | medium | high" reasoning, even when sometimes "medium" wasn't supported and a supported value of "max" was a type error. Now we strictly enforce each model's reasoning separately. This PR is ready for review and merge! Thanks @cjol |
|
The PR looks good to me ✅ |
Refreshes the Workers AI binding types so each reasoning model's types accept exactly the reasoning controls it supports, with per-model hover docs.
This is an intentional TypeScript breaking change for code that passes a reasoning control the model does not support. Those requests were already ignored, remapped or rejected at runtime; now they fail to compile instead.
What changes
reasoning_effort(and Responsesreasoning.effort) accepts exactly the model's supported efforts from Workers AI ConfigAPI. Compatibility aliases (for example GLM-5.2low, which maps tohigh) are not typed. The hover lists them as accepted by the API but not by the types.reasoning_effort. Theirenable_thinkinghover explains that.reasoning_effortenable_thinkingmaxhighnonebooleanmaxhighlownonebooleanmaxhighlowtruehighnonebooleanlowmediumxhighbooleanlowmediumhightruebooleantruechat_template_kwargs)booleanlowmediumhighbooleanCompatibility
Checked against
main'sgenerated-snapshot/index.ts, imported side by side:postProcessedOutputsis identical for all 97 models in both.inputsnarrows for GPT-OSS 20B/120B, Kimi K2.6, Kimi K2.7 Code, GLM-5.2, DeepSeek V4 Flash/Pro, Qwen3.8 and GLM-5.3-Flash. Gemma 4, GLM-4.7-Flash and Nemotron 3 dropreasoning_effort, so object literals that pass it no longer compile. Every other model's inputs are unchanged.AiModelskey removed is img2img.types/test/types/ai-reasoning.tspins each model's exact effort union and covers supported calls, Nemotron's controls, and rejected values: unsupported efforts, aliases, arbitrary strings, the shared effort type on GLM-5.3, effort on models without levels,enable_thinking: falseon mandatory models, and Responsesminimal.Generation
The declarations come from the SDK generator in cloudflare/ai/sdk!1778, run against production ConfigAPI. Only the reasoning, Nemotron and new-model declarations are applied to
types/defines/ai.d.ts, because the rest of that file has drifted from the generator.generated-snapshot/is the output ofbazel build //types, andbazel test //types/...passes (29/29). The branch includes a merge ofupstream/main.The docs counterpart is cloudflare-docs#33541.