BeamWeaver OpenAI
The first OpenAI slices target non-Azure OpenAI paths used by the pinned LangChain OpenAI package.
The Responses, Chat Completions, and embeddings wire contract was compared on September 23, 2026 against
openai-openapi commit 946e365
.
Implemented
-
BeamWeaver.OpenAI.ChatModelimplementsBeamWeaver.Core.ChatModel. -
BeamWeaver.OpenAI.EmbeddingModelimplementsBeamWeaver.Core.EmbeddingModel. -
BeamWeaver.OpenAI.generate_image/2calls the direct Images API and returns the provider response. It traces the request and usage without storing image bytes in the trace:{:ok, response} = BeamWeaver.OpenAI.generate_image("A blue square on white", model: "gpt-image-2.5-flare", quality: "low", size: "1024x1024" ) image = hd(response["data"])["b64_json"] -
Requests go through
BeamWeaver.Transport, so provider tests can run against replay cassettes and live calls can use Req/Finch. -
BeamWeaver messages become Responses API
inputitems. -
BeamWeaver tools become OpenAI
functiontool declarations. -
BeamWeaver.OpenAI.ToolCallingbuilds Responses API built-in tool declarations for web search, file search, code interpreter, image generation, MCP, custom tools, tool search, and OpenAI'sapply_patchtool. -
BeamWeaver.OpenAI.Responsesbuilds raw multi-turn input items and extracts preserved output items from assistant responses. -
Structured output options become
text.formatJSON schema requests. -
Strict structured-output schemas are normalized for OpenAI: object schemas are closed, optional properties become nullable required fields, stale
requiredentries are dropped, and unsupported validation/composition keywords are removed before request rendering. -
Responses API request options include reasoning, include, previous response IDs, raw input items, tool choice, truncation, text verbosity, context management, metadata, service tier, store, model kwargs, and supported sampling controls.
-
Opt-in
include_response_headerspreserves normalized transport response headers on assistant message metadata for sync and Task-backed async chat calls. -
Responses API JSON results become assistant messages with text, function tool calls, and preserved built-in output blocks.
-
Responses and Chat Completions accept the provider's
moderationrequest configuration and expose returned moderation results in response metadata. -
Structured output responses are JSON-decoded into message metadata, and caller parser/validator failures return an OpenAI error with the assistant response attached for debugging.
-
LangChain v3 Responses edge blocks are preserved across input and output: assistant
function_callitems, grouped assistant text/refusal message blocks, invalid function-call argument strings, web/file search outputs, response errors, incomplete details, and provider metadata. -
Multi-turn helpers preserve custom tool calls, image generation calls, MCP approval requests, encrypted reasoning items, and assistant message output items so they can be sent back in the next Responses API turn.
-
Assistant output items replayed through normal messages strip BeamWeaver internal fields such as
raw_provider_block, while reasoning blocks keep only OpenAI-accepted replay fields. -
store: falsereplay sanitization is applied before the request body is sent: provider-only output item IDs are dropped, encrypted reasoning is preserved, non-replayable reasoning is skipped, and empty image-generation placeholders are not sent back to OpenAI. -
Responses API output parsing preserves
apply_patch_callandapply_patch_call_outputitems as provider-scoped content blocks so cached or streamed turns can be replayed without losing patch metadata. -
Current OpenAPI output items such as local/hosted shell calls, programmatic tool calls, computer calls, function outputs, and additional tools are retained as provider-scoped content blocks. Shell lifecycle frames are also exposed as typed custom stream events.
-
Embeddings support document/query calls, dimensions, caller chunk size, Task-backed async calls, and opt-in
skip_emptyhandling. -
Responses API and chat-completions SSE bodies are parsed into text deltas.
-
OpenAI GPT model profiles expose
tool_call_streaming: truewhen the checked-in profile supports incremental streamed tool-call arguments. -
BeamWeaver.OpenAI.Streaming.response/1reconstructs final Responses API output items from SSE streams for text, reasoning summaries, function calls, and terminal built-in tool output items. -
Streaming image generation requests add
partial_images: 1when needed, and partial image frames are preserved on reconstructedimage_generation_calloutput items. -
BeamWeaver.OpenAI.ChatModel.stream_response/3consumes a streaming Responses API response and returns a reconstructed assistant message when callers need tool calls or raw output items, whilestream/3remains the text-chunk API. -
BeamWeaver.OpenAI.ChatModel.async_invoke/3,async_batch/3,async_stream/3,async_stream_response/3, andasync_stream_events/3expose Task-backed async public APIs. Embeddings expose matching Task-backed async invoke and batch helpers. -
BeamWeaver.OpenAI.Streaming.lifecycle_events/1andBeamWeaver.OpenAI.ChatModel.stream_events/3expose message and content-block lifecycle events for streamed text, reasoning, tool calls, and built-in output blocks. -
Chat Completions streams preserve empty initial role-only chunks, incremental tool argument deltas, final assistant tool calls, finish reasons, and detailed usage metadata.
-
OpenAI namespace constructors load defaults from
config :beam_weaver, :openai; put any OS environment reads in yourconfig/runtime.exs. Explicit options still win. Custom routing uses explicit:endpointoptions on the provider model/client structs. -
Streamed tool-search loops preserve
tool_search_call,tool_search_output, streamedfunction_call, and the follow-upfunction_call_outputturn through raw Responses API input items. -
Agent-style function loops preserve raw reasoning/function-call output items and follow-up
function_call_outputinput items in both JSON and streaming Responses API paths. -
Streamed compaction preserves the server
compactionoutput item so later turns can send it back unchanged. -
Phase-tagged output text preserves
phasemetadata for commentary and final answer blocks, including streamed lifecycle events. -
Chat Completions audio input blocks are converted to OpenAI
input_audiocontent parts, and audio output parts are preserved as message content blocks and metadata. The current Responses request schema exposes neitherinput_audiocontent nor top-levelaudioormodalitiescontrols, so current OpenAI Responses profiles do not advertise them. -
Current OpenAI reasoning-model request controls follow provider constraints:
max_tokensandmax_completion_tokensmap to Responsesmax_output_tokens, and temperature is omitted for GPT-5 models unless reasoning effort isnone. Astra rejects temperature and the other unsupported sampling and logprob controls before transport. -
GPT-5.6 requests support
prompt_cache_options, explicit cache breakpoints on content parts, Responses filedetail, normalized cache-write usage, and a model-levelsafety_identifier. -
GPT-5.6 Chat Completions function tools are rejected before transport unless the effective reasoning effort is
none; reasoning plus tools uses Responses. -
GPT-6 Astra and GPT-6.1 Sol support tool-free Chat Completions, but tool calling is rejected there before transport and directed to Responses. Their reasoning effort must be
low,medium,high,xhigh, ormax. -
GPT-6 Sol and Luna support
none,low,medium,high,xhigh, andmaxreasoning. Chat Completions function tools require explicitnoneeffort; reasoning with tools uses Responses. Sampling controls that requirenoneeffort are rejected when reasoning is active. -
Responses accepts
access_programsfor an explicit Cyber access program and preserves the effective selection in response metadata. Nestedprompt_cache_options.prewarmpasses through for cache-only requests. Currentresponse.compaction.compactingevents, optional function-call names, and image-generation output fields are preserved by the existing stream and provider-output paths. Runmix run examples/openai_gpt6_controls.exswith an OpenAI API key to check access-program selection and cache prewarming live. -
Replay-backed provider tests cover the OpenAI cassette shapes that map to BeamWeaver's Responses-oriented chat model.
GPT-6 Astra Profile
BeamWeaver includes gpt-6-astra as a first-class OpenAI profile. It exposes a 1.05M-token context window, 128K maximum output, text and image input, Responses and tool-free Chat Completions, function calling through Responses, structured output, streaming, prompt caching, persisted reasoning, and pro reasoning mode. The profile records OpenAI's current Trusted Access rollout rather than implying general availability.
Current prices per million tokens are $10.00 input, $1.00 cached input, $12.50 cache write, and $50.00 output. Requests above 272K input tokens use 2x input and cache rates and 1.5x output for the full request. Batch and Flex cost half the Standard rate, while Fast costs twice the applicable rate. Fast is not available for Astra with EU data residency. See the GPT-6 Astra model page .
BeamWeaver.Models.init_chat_model!("openai:gpt-6-astra",
reasoning: %{effort: :high},
prompt_cache_options: %{mode: :explicit, ttl: "30m"}
)
Astra does not accept none or minimal reasoning, temperature, top_p, or top_logprobs; Chat Completions also rejects logprobs. Responses include cannot request message.output_text.logprobs.
GPT-6.1 Sol Profile
BeamWeaver includes openai:gpt-6.1-sol for complex coding, computer use, and professional work. The model accepts text and image input and produces text, with a 1,050,000-token context window, 922,000-token maximum input, and 128,000-token maximum output. It supports streaming, structured output, prompt caching, and function and provider-hosted tools through Responses. Chat Completions is available for requests without tools.
Reasoning efforts are low, medium (default), high, xhigh, and max. none, minimal, temperature, top_p, and logprob controls are rejected before transport. The default API is Responses.
Standard pricing per million tokens is:
| Input context | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| Up to 272K input tokens | $2.00 | $0.10 | $2.50 | $10.00 |
| Above 272K input tokens | $4.00 | $0.20 | $5.00 | $15.00 |
Batch and Flex cost half the applicable Standard rates; Fast costs twice those rates. Eligible regional processing adds 10%. US and EU data residency are supported; Fast is unavailable with EU data residency. See the model specification and pricing .
BeamWeaver.Models.init_chat_model!("openai:gpt-6.1-sol",
reasoning: %{effort: :high},
prompt_cache_options: %{mode: :explicit, ttl: "30m"}
)
OpenAI also offers hosted Multi-agent in beta . The profile records that provider capability. BeamWeaver's local subagents use its existing agent harness; hosted Multi-agent orchestration and agent-specific output handling do not have a dedicated wrapper.
GPT-6 Sol and Luna Profiles
BeamWeaver includes gpt-6-sol and gpt-6-luna as first-class OpenAI profiles. Both support text and image input, text output, streaming, structured output, 1.05M-token context windows, and 128K maximum output. Responses supports reasoning with function tools; Chat Completions function tools require reasoning_effort: :none. The default reasoning effort is medium.
Standard short-context prices per million tokens are:
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
gpt-6-sol
| $2.00 | $0.20 | $2.50 | $10.00 |
gpt-6-luna
| $0.10 | $0.01 | $0.125 | $0.50 |
Above 272K input tokens, input and cache rates double and output rates rise by 50% for the full request. Batch and Flex cost half the Standard rate; Fast costs twice the applicable rate. EU data residency is available only with Standard processing. See the Sol , Luna , and pricing pages.
BeamWeaver.Models.init_chat_model!("openai:gpt-6-sol",
reasoning: %{effort: :high},
prompt_cache_options: %{mode: :explicit, ttl: "30m"}
)
GPT-5.6 Profiles
BeamWeaver includes gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna as first-class OpenAI profiles. The official gpt-5.6 alias resolves to Sol. All three profiles expose a 1.05M-token context window, 128K maximum output, text and image input, Responses and Chat Completions, function calling, structured output, streaming, and the current OpenAI built-in tool catalog.
Current short-context prices per million tokens are:
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
gpt-5.6-sol
| $4.00 | $0.40 | $5.00 | $20.00 |
gpt-5.6-terra
| $2.00 | $0.20 | $2.50 | $12.00 |
gpt-5.6-luna
| $0.20 | $0.02 | $0.25 | $1.20 |
The Sol row is promotional through November 21, 2026. Its checked-in regular rates remain $5.00 input, $0.50 cached input, and $30.00 output per million tokens; the corresponding regular cache-write rate is $6.25.
Requests above 272K input tokens use OpenAI's higher-context rates: 2x input and 1.5x output for the full request. GPT-5.6 cache writes cost 1.25x uncached input, cache reads receive the 90% discount, and the current cache TTL is 30 minutes. Eligible regional-processing endpoints add OpenAI's 10% uplift. Batch and Flex processing cost half the Standard rate. Fast mode costs twice the Standard short-context rate and replaces Priority Processing; OpenAI continues to accept both service_tier: :fast and service_tier: :priority. See
OpenAI API pricing
.
The existing reasoning request option carries GPT-5.6 controls without a separate model type:
BeamWeaver.Models.init_chat_model!("openai:gpt-5.6-sol",
reasoning: %{effort: :max, mode: :pro, context: :all_turns}
)
Persisted reasoning context returned by Responses is preserved as message.response_metadata.reasoning_context. Explicit GPT-5.6 prompt caching is available on both OpenAI APIs:
BeamWeaver.Models.init_chat_model!("openai:gpt-5.6",
prompt_cache_options: %{mode: :explicit, ttl: "30m"},
safety_identifier: "user_<stable_privacy_preserving_hash>"
)
Mark the desired text, image, or file content block with metadata.prompt_cache_breakpoint. Cache reads and writes are normalized to usage_metadata.input_token_details.cache_read and cache_write. Supplying comparison_response_id inside prompt_cache_options is passed through, and returned diagnostics are available at message.response_metadata.prompt_cache_diagnostics.
On Chat Completions, GPT-5.6 function tools require reasoning_effort: :none. BeamWeaver returns :invalid_model_option before transport when this combination is invalid and points callers to Responses.
OpenAI's programmatic tool calling and hosted multi-agent beta are recorded as provider capabilities in profile metadata. BeamWeaver does not yet provide dedicated request/response helpers for those two hosted orchestration surfaces.
Codex-Flavoured Responses Contracts
BeamWeaver.OpenAI.CodexResponses contains pure helpers for an embedding application that has already admitted a Codex-flavoured Responses endpoint and authentication principal. It does not obtain, refresh, persist, or authorize an access token and does not choose an endpoint.
headers = BeamWeaver.OpenAI.CodexResponses.headers(access_token)
request_opts =
BeamWeaver.OpenAI.CodexResponses.request_options(
reasoning: %{effort: "high", summary: "auto"}
)
{:ok, body} =
BeamWeaver.OpenAI.CodexResponses.normalize_body(rendered_body,
maximum_body_bytes: 8_388_608
)
The helpers enforce store: false, request encrypted reasoning content by default, remove unsupported output/tool-count limits, avoid an empty tools array, and canonicalize controlled atom/string keys without emitting duplicate wire parameters. headers/2 emits only the compatibility headers and, when the JWT claim is present, the ChatGPT account ID; bearer authorization remains the caller's responsibility. Invalid or oversized bodies fail before transport.
Replay Usage
model =
BeamWeaver.OpenAI.chat_model(
api_key: "sk-replay",
transport: BeamWeaver.Transport.Replay,
transport_opts: [cassette_path: "path/to/my_agent_response.yaml"]
)
BeamWeaver.Core.ChatModel.invoke(model, [
BeamWeaver.Core.Message.user("agent ping")
])
Built-in tool declarations are plain request values:
tools = [
BeamWeaver.OpenAI.ToolCalling.web_search(),
BeamWeaver.OpenAI.ToolCalling.file_search(["vs_123"]),
BeamWeaver.OpenAI.ToolCalling.code_interpreter(%{type: :auto}),
BeamWeaver.OpenAI.ToolCalling.apply_patch()
]
Multi-turn output items can be fed back explicitly:
{:ok, first} = BeamWeaver.Core.ChatModel.invoke(model, messages, tools: tools)
input_items =
[BeamWeaver.OpenAI.Responses.message(:user, "original prompt")] ++
BeamWeaver.OpenAI.Responses.output_items(first) ++
[BeamWeaver.OpenAI.Responses.custom_tool_call_output("call_123", "27")]
BeamWeaver.Core.ChatModel.invoke(model, [], input_items: input_items, tools: tools)
The replay matcher compares method, URL, and canonical JSON body. Tests for tools, structured output, reasoning, MCP approvals, and tool search intentionally depend on that matcher so request shape regressions fail before a live provider call is involved.
For cached multi-turn Responses replay, pass store: false through model options or extra_body. BeamWeaver will keep the replayable parts of assistant history while removing provider-generated item IDs that OpenAI rejects when storage is disabled:
BeamWeaver.Core.ChatModel.invoke(model, messages, extra_body: %{store: false})
Replay coverage for this request shape includes encrypted reasoning preservation and provider-only ID removal.
Run the supervised demo with:
mix run examples/supervised_agent.exs
Inspect the OpenAI apply_patch built-in tool request shape without live credentials:
mix run examples/openai_apply_patch_tool.exs
Unsupported OpenAI Surfaces
-
Azure OpenAI.
-
ChatOpenAI,OpenAIEmbeddings, andOpenAIPython compatibility surfaces that do not map to the current BeamWeaver chat, embeddings, and Responses APIs. -
Callback/client wrapper behavior outside the Task-backed async public APIs.
-
Audio API surfaces beyond the current chat audio request/response shape.
-
GPT-5.6/Astra programmatic tool calling, Astra async tool calling and mid-turn steering, and OpenAI-hosted multi-agent orchestration.
-
LangChain v3 protocol edge cases outside the current Responses message conversion layer.
-
WeaveScope exporter behavior beyond the native BeamWeaver trace payload boundary.