English
API Changes – V12.0
Added
- Multimodality / input assets:
POST /v1/input-asset/upload/init– Initialise an input asset upload (presigned S3 URL,input_asset_id).POST /v1/input-asset/upload/complete– Complete the upload and start processing.DELETE /v1/input-asset/upload/abort– Abort an active upload and clean up the resource.PATCH /v1/chat-session/input-assets/– Bulk update input assets.DELETE /v1/chat-session/input-assets/– Bulk delete input assets.PATCH /v1/chat-session/input-asset/{asset_id}– Update a single input asset.DELETE /v1/chat-session/input-asset/{asset_id}– Delete a single input asset.
- Reasoning control:
- New
reasoning_effortparameter inStreamRequest,ChatSession,NewChatSessionSchemaandPromptTemplatefor client-side control of thinking intensity on supported models (OpenAI o-series, Anthropic Claude 3.7+, Gemini).
- New
- Transcriptions (new S3 workflow):
- New three-step upload workflow for transcriptions using S3 presigned URLs (supports files up to 500 MB):
POST /v1/transcript/upload/init– Initialises the upload. Returnstranscript_id,upload_mode(singleormultipart), presigned URL(s), and, where applicable,upload_idandpart_urlsfor multipart uploads.PUT <presigned_url>– Uploads the file (or individual parts) directly to S3 without passing through the API server (URL from the initialisation step).POST /v1/transcript/upload/complete– Completes the upload and starts transcription. For multipart uploads,upload_idandparts(ETags) are required.DELETE /v1/transcript/upload/abort– Aborts an active multipart upload and cleans up the transcript resource. Requirestranscript_idandupload_id.
- New three-step upload workflow for transcriptions using S3 presigned URLs (supports files up to 500 MB):
Changed
- Chat streaming (
/v1/chat/stream/{session_id}):- Multimodal support: The endpoint now automatically detects unassigned input assets in the session and includes them in the AI request.
- Format change: The response now uses NDJSON (
application/x-ndjson) instead of Server-Sent Events (SSE). - Reasoning: Supports the
reasoning_effortparameter for dynamic control of model thinking processes.
- Chat providers and reasoning: Replaced static provider variants with a dynamic reasoning parameter (see details).
- Input asset management:
/v1/chat-session/{session_id}/sequence/{sequence_id}/input-asset/{asset_id}now supportsPATCH(update) andDELETE(delete), in addition to the simplified/v1/chat-session/input-asset/{asset_id}path.
Deprecated
POST /v1/transcript/– Direct file uploads for transcriptions are deprecated and will be removed in a future version.⚠ Breaking change:
- The limit is reduced to 50 MB. The new upload workflow is then mandatory for all files larger than 50 MB (see migration).
Early migration to the new workflow is strongly recommended.
Responses from this endpoint include the following HTTP header according to RFC 7234:
Warning: 299 - "This API call is deprecated and will be removed. Refer release notes for details."Migrating to the new upload workflow:
Replace the previous direct upload with the following three-step workflow:
Step 1 – Initialise the upload:
httpPOST /v1/transcript/upload/init Content-Type: application/json Authorization: Bearer <token> { "filename": "meeting.mp4", "file_size": 52428800, "language": "de-DE", "diarization": true }Response:
json{ "transcript_id": "abc123", "upload_mode": "single", "file_path": "abc123-meeting.mp4", "upload_url": "https://s3.amazonaws.com/bucket/meeting.mp4?X-Amz-Signature=...", "upload_id": null, "part_urls": null, "part_size": null, "expires_in_seconds": 3600 }Step 2 – Upload the file directly to S3:
The
upload_modefrom the initialisation response determines the upload path (singleormultipart).http### Single part (upload_mode: "single") PUT https://s3.amazonaws.com/bucket/meeting.mp4?X-Amz-Signature=... Content-Type: video/mp4 <binary file content> ### Multipart (upload_mode: "multipart") – repeat for each part PUT https://s3.amazonaws.com/bucket/meeting.mp4?partNumber=1&uploadId=...&X-Amz-Signature=... Content-Type: video/mp4 <binary chunk 1> # Store the ETag from the response header → pass it to the complete callpythonimport requests # Values from the initialisation response if upload_mode == "single": with open("meeting.mp4", "rb") as f: response = requests.put( upload_url, data=f, headers={"Content-Type": "video/mp4"}, ) response.raise_for_status() parts = None elif upload_mode == "multipart": parts = [] with open("meeting.mp4", "rb") as f: for i, part_url in enumerate(part_urls, start=1): chunk = f.read(part_size) response = requests.put( part_url, data=chunk, headers={"Content-Type": "video/mp4"}, ) response.raise_for_status() parts.append({"part_number": i, "etag": response.headers["ETag"]}) # Pass parts to the complete call (see step 3)javascriptimport { createReadStream, statSync } from "fs"; // Values from the initialisation response let parts = null; if (uploadMode === "single") { const response = await fetch(uploadUrl, { method: "PUT", body: createReadStream("meeting.mp4"), headers: { "Content-Type": "video/mp4", "Content-Length": String(statSync("meeting.mp4").size), }, }); if (!response.ok) throw new Error(`S3 upload failed: ${response.status}`); } else if (uploadMode === "multipart") { parts = []; const fileSize = statSync("meeting.mp4").size; let offset = 0; for (let i = 0; i < partUrls.length; i++) { const chunkSize = Math.min(partSize, fileSize - offset); const response = await fetch(partUrls[i], { method: "PUT", body: createReadStream("meeting.mp4", { start: offset, end: offset + chunkSize - 1 }), headers: { "Content-Type": "video/mp4", "Content-Length": String(chunkSize), }, }); if (!response.ok) throw new Error(`Part ${i + 1} failed: ${response.status}`); parts.push({ part_number: i + 1, etag: response.headers.get("ETag") }); offset += chunkSize; } } // Pass parts to the complete call (see step 3)No
Authorizationheader is required—the authentication is already included in the presigned URL.Step 3 – Complete the upload and start processing:
Single part:
httpPOST /v1/transcript/upload/complete Content-Type: application/json Authorization: Bearer <token> { "transcript_id": "abc123", "upload_id": null, "parts": null }Multipart (ETags from step 2 required):
httpPOST /v1/transcript/upload/complete Content-Type: application/json Authorization: Bearer <token> { "transcript_id": "abc123", "upload_id": "some-upload-id", "parts": [ { "part_number": 1, "etag": "\"abc123\"" }, { "part_number": 2, "etag": "\"def456\"" }, { "part_number": 3, "etag": "\"ghi789\"" } ] }Check status (unchanged):
httpGET /v1/transcript/abc123 Authorization: Bearer <token>You can retrieve the transcription status with
GET /v1/transcript/{transcript_id}(possible values:new,succeeded,failed).
Removed
- Instruct API:
- The
POST /v1/instruct/httpandPOST /v1/instruct/streamendpoints have been removed. Please use the corresponding/v1/chatendpoints instead.
- The
- Specific reasoning providers:
- Provider variants with fixed reasoning levels in their names (for example,
azure-gpt5-high-reasoning-effortorazure-openai-gpt-5_1-low-reasoning-effort) have been removed. - Use the base provider (for example,
azure-openai-gpt-5_4) with thereasoning_effortparameter instead.
- Provider variants with fixed reasoning levels in their names (for example,
Chat providers and reasoning effort
With version 12.0, reasoning intensity (thinking) selection has moved from the provider level to the parameter level. This enables more flexible control without changing the provider.
The reasoning_effort parameter
The parameter can be set when creating a chat session or directly in the streaming request (POST /v1/chat/stream/{session_id}).
Permitted values:
| Value | UI label | Description |
|---|---|---|
none | Off | Disables reasoning or extended-thinking features. |
low | Standard | Uses a basic reasoning level or moderate thinking budget. |
high | Verbose | Uses a higher reasoning level or larger thinking budget. |
How it works per provider:
The API abstracts provider-specific implementations:
- OpenAI (o-series): Maps directly to the native
reasoning_effortfield. - Anthropic (Claude 3.7+): Controls the
thinkingbudget (for example, 1024 versus 4096 tokens). - Google Gemini: Controls
thinking_budget(tokens) orthinking_level.