OpenAI Clients¶
Configure an ordinary OpenAI SDK with the LGOS base URL, usually
http://localhost:8000/v1, and use a registered graph name as model. The
api_key is sent as a bearer token; the application from
Get Started does not verify it.
Use Responses for new clients and every maintained demo UI. Chat Completions
remains a direct compatibility surface for clients that cannot use Responses.
An optional proxy must expose a native /v1/responses route and preserve the
same typed items and events; raw pass-through is not part of the client design.
The basic examples call the provider-free echo graph from the getting-started
application. Names such as research-graph and interruptible describe
capabilities the registered graph must enable. The
demo graph catalog provides runnable examples.
Install A Client¶
Create A Response¶
LGOS is stateless at the Responses layer. Send store=False explicitly so the
request remains portable to other OpenAI-compatible endpoints; LGOS also
defaults an omitted value to false.
Do not expose real API keys in a browser
The JavaScript example enables dangerouslyAllowBrowser only because the
local application uses a dummy key. Keep production credentials in
server-side code.
Stream Final Text And Commentary¶
Responses streams typed lifecycle events rather than Chat chunks. When a graph
declares GraphFeature.CLIENT_EVENTS, every visible status_event() becomes a
completed assistant message with phase="commentary"; no request metadata
opt-in is required. The durable answer uses phase="final_answer".
Track the phase from response.output_item.added before handling text deltas:
phases = {}
with client.responses.stream(
model="research-graph",
input="Research this topic.",
store=False,
) as stream:
for event in stream:
if (
event.type == "response.output_item.added"
and event.item.type == "message"
):
phases[event.output_index] = event.item.phase
elif event.type == "response.output_text.delta":
if phases.get(event.output_index) == "final_answer":
print(event.delta, end="", flush=True)
elif event.type == "response.output_text.done":
if phases.get(event.output_index) == "commentary":
show_status(event.text)
response = stream.get_final_response()
const phases = new Map();
const stream = await openai.responses.create({
model: "research-graph",
input: "Research this topic.",
store: false,
stream: true,
});
for await (const event of stream) {
if (
event.type === "response.output_item.added" &&
event.item.type === "message"
) {
phases.set(event.output_index, event.item.phase);
} else if (
event.type === "response.output_text.delta" &&
phases.get(event.output_index) === "final_answer"
) {
process.stdout.write(event.delta);
} else if (
event.type === "response.output_text.done" &&
phases.get(event.output_index) === "commentary"
) {
showStatus(event.text);
}
}
Commentary is streaming-only and transient. Non-streaming execution does not
collect old status updates. The OpenAI Python SDK's Response.output_text
convenience property concatenates all output-text parts, including commentary,
so a UI processing a completed stream must select only message items whose
phase is final_answer:
final_text = "".join(
part.text
for item in response.output
if item.type == "message" and item.phase == "final_answer"
for part in item.content
if part.type == "output_text"
)
The maintained Chainlit and Open WebUI adapters apply this filter and render
commentary through their native status interfaces. A generic client that does
not understand phase may display commentary as answer text; keep status text
useful but do not treat that client as a full advanced-UI integration.
Manage Conversation State¶
LGOS does not persist Response objects or Conversations. It rejects
store=True, conversation, and background mode (previous_response_id is
supported only for resuming interruptible graphs).
Keep an input ledger and resend the items needed by each turn:
input_items = [{"role": "user", "content": "Introduce LangGraph briefly."}]
first = client.responses.create(
model="echo",
input=input_items,
store=False,
)
input_items.extend(item.model_dump(mode="json") for item in first.output)
input_items.append({"role": "user", "content": "Now make it one sentence."})
second = client.responses.create(
model="echo",
input=input_items,
store=False,
)
Replay complete SDK output items instead of rebuilding assistant text. This
preserves item IDs, function-call IDs, and assistant phase. Keep every earlier
user, system, or developer item that the next turn needs. This is application
conversation state; LangGraph checkpoints remain a separate temporary store for
paused interrupts.
Continue Function Calls¶
LGOS accepts the flat Responses function-tool shape. When a graph returns a
function_call, execute only a function your client owns, then replay the
complete output and append the matching string-valued result:
import json
tools = [
{
"type": "function",
"name": "lookup_order",
"description": "Look up an order visible to the signed-in user.",
"strict": True,
"parameters": {
"type": "object",
"additionalProperties": False,
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
}
]
input_items = [{"role": "user", "content": "Where is order A123?"}]
response = client.responses.create(
model="my-graph",
input=input_items,
tools=tools,
store=False,
)
input_items.extend(item.model_dump(mode="json") for item in response.output)
for item in response.output:
if item.type != "function_call":
continue
result = lookup_order(**json.loads(item.arguments))
input_items.append(
{
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(result),
}
)
completed = client.responses.create(
model="my-graph",
input=input_items,
tools=tools,
store=False,
)
The same full-item rule drives the demo's Files-plus-display_file chart
contract. The trusted UI downloads the file_id, renders or persists it
natively, and returns a small acknowledgment. See
Accept And Display Files.
Resume An Interrupt¶
An interrupt-enabled graph returns one or more function_call items named
lgos_interrupt. Preserve every returned call and answer the whole batch.
No metadata is required for an initial request; use a new UUID in
metadata.lgos_run_id when retrying a lost initial response must address
the same pending operation.
import json
from uuid import uuid4
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="DUMMY",
max_retries=0,
)
metadata = {"lgos_run_id": str(uuid4())}
input_items = [
{"role": "user", "content": "Perform the protected action."}
]
paused = client.responses.create(
model="interruptible",
input=input_items,
metadata=metadata,
store=False,
)
calls = [item for item in paused.output if item.type == "function_call"]
if not calls:
raise RuntimeError("The graph completed without interrupting")
# Resume using standard previous_response_id:
completed = client.responses.create(
model="interruptible",
previous_response_id=paused.id,
input=[
{
"type": "function_call_output",
"call_id": call.call_id,
"output": collect_answer(call),
}
for call in calls
],
metadata=metadata,
store=False,
)
Do not resume a subset, synthesize a new call, or send only the visible
question. Persist paused.id and the returned calls before asking the user so
a reconnect can reproduce the request. Runtime settings remain per-request and
must be resent. See
Tool Calls And Interrupts
for stale-state conflicts and recovery boundaries.
Model Discovery And Runtime Settings¶
Use client.models.list() for registered graph IDs. Direct LGOS model objects
also expose the namespaced lgos extension. Retrieve a
selected model to discover its settings descriptor:
model = client.models.retrieve("my-settings-graph")
extension = (model.model_extra or {}).get("lgos")
settings = (
extension.get("client_settings")
if isinstance(extension, dict) and extension.get("schema_version") == 1
else None
)
if isinstance(settings, dict) and settings.get("schema_version") == 1:
print(settings["json_schema"])
print(settings["defaults"])
metadata.lgos_settings is JSON text, not a nested metadata
object. Send only values that differ from the advertised defaults, keep the
encoded value at 512 characters or fewer, and resend it on every request that
needs it. A normalizing proxy may omit this optional extension; plain Responses
still work, but the client must not infer settings or graph features it cannot
discover. See Runtime Settings.
Direct Chat Compatibility¶
Chat Completions remains available for direct clients that need the familiar message/choice shape:
completion = client.chat.completions.create(
model="echo",
messages=[{"role": "user", "content": "Hello through Chat"}],
)
print(completion.choices[0].message.content)
This route shares the same graph runner but has its own protocol adapter. Chat
Completions is suited for simple graphs and tool calls. For advanced workflows
such as streaming status commentary, checkpointed persistence, or interrupts,
use the Responses API (/v1/responses).
Diagnostics¶
Direct HTTP diagnostic
Use direct HTTP only to inspect behavior while debugging:
Set timeouts for long-running graphs and add bearer-token authentication before exposing the API outside a trusted development environment. The exact accepted request fields and explicit exclusions are in the compatibility contract.