OpenAI Clients¶
Configure clients with the server base URL, usually http://localhost:8000/v1.
The api_key value is sent as Authorization: Bearer <key>; the application
from Get Started does not verify it.
For one LGOS deployment, use one OpenAI client and base URL for model listing, model retrieval, and chat completions. When a proxy is present, that URL must be its complete LGOS pass-through route. A gateway that federates multiple LGOS providers may expose a separate routing catalog; use it only to discover provider-qualified IDs, then use pass-through for LGOS metadata and chat. The Bifrost demo shows that split.
The basic examples below call that application's echo model. Examples named
my-graph, my-settings-graph, or research-graph describe capabilities your
registered graph must enable. The demo graph catalog
provides runnable models for those advanced behaviors.
Install A Client¶
Chat Completions¶
Do not expose real API keys in a browser
The JavaScript examples enable dangerouslyAllowBrowser because the local
examples use a dummy key. Keep production credentials in server-side code.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="DUMMY")
stream = client.chat.completions.create(
model="my-graph",
messages=[{"role": "user", "content": "Write a short poem about graphs."}],
stream=True,
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
print(content, end="")
import OpenAI from "openai";
const openai = new OpenAI({
baseURL: "http://localhost:8000/v1",
apiKey: "DUMMY",
dangerouslyAllowBrowser: true,
});
const completion = await openai.chat.completions.create({
model: "echo",
messages: [{ role: "user", content: "Hello from JavaScript" }],
});
console.log(completion.choices[0].message.content);
import OpenAI from "openai";
const openai = new OpenAI({
baseURL: "http://localhost:8000/v1",
apiKey: "DUMMY",
dangerouslyAllowBrowser: true,
});
const stream = await openai.chat.completions.create({
model: "my-graph",
messages: [{ role: "user", content: "Write a short poem about graphs." }],
stream: true,
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content || "";
process.stdout.write(content);
}
Client Stream Events¶
First retrieve the model and confirm that
langgraph_openai_serve.features contains client_events. Then request
explicitly public graph events with the standard metadata field. The Python SDK
keeps the namespaced extension in each chunk's model_extra:
model = client.models.retrieve("research-graph")
model_extension = (model.model_extra or {}).get("langgraph_openai_serve")
features = (
model_extension.get("features") if isinstance(model_extension, dict) else None
)
valid_extension = (
isinstance(model_extension, dict)
and model_extension.get("schema_version") == 1
and isinstance(features, list)
and all(isinstance(feature, str) for feature in features)
)
if not valid_extension:
show_limited_functionality_warning()
event_metadata = (
{"langgraph_stream_events": "v1"}
if valid_extension and "client_events" in features
else {}
)
stream = client.chat.completions.create(
model="research-graph",
messages=[{"role": "user", "content": "Research this topic."}],
stream=True,
metadata=event_metadata,
)
for chunk in stream:
extension = (chunk.model_extra or {}).get("langgraph_openai_serve")
if isinstance(extension, dict) and extension.get("schema_version") == 1:
event = extension.get("event")
if isinstance(event, dict):
handle_client_event(event)
for choice in chunk.choices:
if choice.delta.content:
print(choice.delta.content, end="")
When using the higher-level streaming helper, inspect its raw ChunkEvent:
with client.chat.completions.stream(
model="research-graph",
messages=[{"role": "user", "content": "Research this topic."}],
metadata=event_metadata,
) as stream:
for item in stream:
if item.type != "chunk":
continue
extension = (item.chunk.model_extra or {}).get(
"langgraph_openai_serve"
)
if isinstance(extension, dict) and extension.get("schema_version") == 1:
event = extension.get("event")
if isinstance(event, dict):
handle_client_event(event)
The helper emits a raw chunk event for every Chat Completions chunk. Consume
LGOS events during iteration; do not expect get_final_completion() to retain
them. See the OpenAI SDK's
Chat Completions event reference
and the LGOS wire contract.
Portable status events contain a user-facing description plus done and
hidden booleans. Treat them as passive UI updates, and stop the active
indicator when done is true. Do not execute them as OpenAI tool calls; the
backend graph owns the work.
Model Discovery And Runtime Settings¶
List model summaries and read descriptions from their lightweight LGOS
extensions, then retrieve the selected model to discover its settings. Check
both the LGOS extension version and the nested runtime-settings version. A valid
detail extension without client_settings means the model has no public
settings. A missing description or an invalid detail extension means the
configured endpoint is degraded: keep standard chat available, omit extended
behavior, and show Limited functionality rather than silently treating the
model as fully capable.
models = client.models.list()
selected = next(
model for model in models.data if model.id == "my-settings-graph"
)
summary_extension = (selected.model_extra or {}).get(
"langgraph_openai_serve"
)
description = (
summary_extension.get("description")
if isinstance(summary_extension, dict)
and summary_extension.get("schema_version") == 1
else None
)
print(description)
model = client.models.retrieve(selected.id)
extension = (model.model_extra or {}).get("langgraph_openai_serve")
settings = (
extension.get("client_settings")
if isinstance(extension, dict) and extension.get("schema_version") == 1
else None
)
if isinstance(settings, dict) and settings.get("schema_version") == 1:
print(settings["json_schema"])
print(settings["defaults"])
const models = await openai.models.list();
const selectedModel = models.data.find(
(model) => model.id === "my-settings-graph",
);
if (!selectedModel) throw new Error("my-settings-graph is not registered");
const summaryExtension = selectedModel.langgraph_openai_serve;
const description =
summaryExtension?.schema_version === 1
? summaryExtension.description
: undefined;
console.log(description);
const model = await openai.models.retrieve(selectedModel.id);
const extension = model.langgraph_openai_serve;
const settings =
extension?.schema_version === 1 &&
extension.client_settings?.schema_version === 1
? extension.client_settings
: undefined;
if (settings) {
console.log(settings.json_schema);
console.log(settings.defaults);
}
metadata.langgraph_runtime_settings must be a JSON-encoded string, produced by
json.dumps() or JSON.stringify(), rather than a nested metadata object. Send
only values that differ from the discovered defaults; the encoded value must be
512 characters or fewer. See
Client Request
for the request shape.
Settings apply to one request. Resend non-default values whenever they are needed, including interrupt-resume requests. Omitting the metadata on a later request uses server defaults again.
Interrupt Resume¶
Interrupt-enabled graphs use OpenAI tool calls. Retrieve the selected model and
check langgraph_openai_serve.features for interrupts before starting. Pass
metadata.langgraph_thread_id, then resume langgraph_interrupt with a matching
tool message. The thread ID restores checkpoint state only; include the same
non-default langgraph_runtime_settings string on the resume request when the resumed run
needs those settings. See
OpenAI compatibility.
Diagnostics¶
Direct HTTP diagnostic
Use direct HTTP only to inspect behavior while debugging:
Notes¶
- Use the registered graph name as
model. - Set timeouts for long-running graphs.
- Use streaming only for graphs configured to emit streamed chunks.
- Add bearer-token authentication before exposing the API outside trusted development environments.