Skip to content

Queued Requests (Queue API)

The /queue endpoints work like the /v2 endpoints, with one difference: when you are over your rate limit, /v2 returns 429 Too Many Requests right away, while /queue waits until your limit frees up and then runs the request.

  • Within your limit: the request runs immediately.
  • Over your limit: the request waits. Limits reset every minute, so it is checked again at the start of each minute.
  • There is no maximum wait. The request keeps waiting until it fits or until it times out (your client’s timeout or a network timeout). Set your client timeout to cover the wait plus the response time.

If you close the connection while waiting, the request is cancelled.

A single request with more input tokens than your per-minute token limit never fits, so it waits until it times out. If your requests often wait more than a minute or two, you are sending more than your limit allows: slow down, or ask for a higher limit.

  • Use /queue for batch jobs that sometimes go over the limit, so short bursts wait instead of failing.
  • Use /v2 for interactive apps, where an immediate 429 is better than waiting, and for long-running requests (use streaming).

Every standard endpoint has a /queue equivalent:

Standard EndpointQueue EndpointDescription
POST /v2/chat/completionsPOST /queue/chat/completionsChat completion
POST /v2/completionsPOST /queue/completionsText completion
POST /v2/embeddingsPOST /queue/embeddingsEmbeddings
POST /v2/audio/transcriptionsPOST /queue/audio/transcriptionsAudio transcription
POST /v2/audio/translationsPOST /queue/audio/translationsAudio translation
POST /v2/images/generationsPOST /queue/images/generationsImage generation
POST /v2/images/editsPOST /queue/images/editsImage editing
GET /v2/modelsGET /queue/modelsList models

The request body for each queue endpoint is identical to its standard counterpart — just change the base path from /v2 to /queue.

Submit a chat completion to the queue:

Terminal window
curl -X POST "https://llm-server.llmhub.t-systems.net/queue/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Write a detailed analysis of renewable energy trends in Europe."}
],
"max_tokens": 2000
}'

Process multiple prompts efficiently using the queue:

import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI(
base_url="https://llm-server.llmhub.t-systems.net/queue",
)
prompts = [
"Summarize the benefits of solar energy.",
"Explain how wind turbines generate electricity.",
"Describe the future of hydrogen fuel cells.",
]
async def process_prompt(prompt):
response = await client.chat.completions.create(
model="gpt-oss-120b",
messages=[{"role": "user", "content": prompt}],
max_tokens=500,
)
return response.choices[0].message.content
async def main():
tasks = [process_prompt(p) for p in prompts]
results = await asyncio.gather(*tasks)
for prompt, result in zip(prompts, results):
print(f"Prompt: {prompt[:50]}...")
print(f"Result: {result[:100]}...\n")
asyncio.run(main())
from openai import OpenAI
client = OpenAI(
base_url="https://llm-server.llmhub.t-systems.net/queue",
)
response = client.embeddings.create(
model="text-embedding-bge-m3",
input="The benefits of renewable energy in Europe",
)
print(f"Embedding dimension: {len(response.data[0].embedding)}")
from openai import OpenAI
client = OpenAI(
base_url="https://llm-server.llmhub.t-systems.net/queue",
)
with open("meeting_recording.mp3", "rb") as audio_file:
transcript = client.audio.transcriptions.create(
model="whisper-large-v3",
file=audio_file,
)
print(transcript.text)
from openai import OpenAI
client = OpenAI(
base_url="https://llm-server.llmhub.t-systems.net/queue",
)
result = client.images.generate(
model="gpt-image-2",
prompt="A futuristic data center powered by renewable energy",
)
print(result.data[0].b64_json[:50] + "...")