Queued Requests (Queue API)
The /queue endpoints work like the /v2 endpoints, with one difference: when you are over your rate limit, /v2 returns 429 Too Many Requests right away, while /queue waits until your limit frees up and then runs the request.
How It Works
Section titled “How It Works”- Within your limit: the request runs immediately.
- Over your limit: the request waits. Limits reset every minute, so it is checked again at the start of each minute.
- There is no maximum wait. The request keeps waiting until it fits or until it times out (your client’s timeout or a network timeout). Set your client timeout to cover the wait plus the response time.
If you close the connection while waiting, the request is cancelled.
A single request with more input tokens than your per-minute token limit never fits, so it waits until it times out. If your requests often wait more than a minute or two, you are sending more than your limit allows: slow down, or ask for a higher limit.
When to Use the Queue API
Section titled “When to Use the Queue API”- Use
/queuefor batch jobs that sometimes go over the limit, so short bursts wait instead of failing. - Use
/v2for interactive apps, where an immediate429is better than waiting, and for long-running requests (use streaming).
Endpoint Mapping
Section titled “Endpoint Mapping”Every standard endpoint has a /queue equivalent:
| Standard Endpoint | Queue Endpoint | Description |
|---|---|---|
POST /v2/chat/completions | POST /queue/chat/completions | Chat completion |
POST /v2/completions | POST /queue/completions | Text completion |
POST /v2/embeddings | POST /queue/embeddings | Embeddings |
POST /v2/audio/transcriptions | POST /queue/audio/transcriptions | Audio transcription |
POST /v2/audio/translations | POST /queue/audio/translations | Audio translation |
POST /v2/images/generations | POST /queue/images/generations | Image generation |
POST /v2/images/edits | POST /queue/images/edits | Image editing |
GET /v2/models | GET /queue/models | List models |
The request body for each queue endpoint is identical to its standard counterpart — just change the base path from /v2 to /queue.
Basic Usage
Section titled “Basic Usage”Submit a chat completion to the queue:
curl -X POST "https://llm-server.llmhub.t-systems.net/queue/chat/completions" \ -H "Authorization: Bearer $OPENAI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-oss-120b", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Write a detailed analysis of renewable energy trends in Europe."} ], "max_tokens": 2000 }'from openai import OpenAI
client = OpenAI( base_url="https://llm-server.llmhub.t-systems.net/queue",)
response = client.chat.completions.create( model="gpt-oss-120b", messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Write a detailed analysis of renewable energy trends in Europe."}, ], max_tokens=2000,)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://llm-server.llmhub.t-systems.net/queue",});
const response = await client.chat.completions.create({ model: "gpt-oss-120b", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: "Write a detailed analysis of renewable energy trends in Europe." }, ], max_tokens: 2000,});
console.log(response.choices[0].message.content);Batch Processing Example
Section titled “Batch Processing Example”Process multiple prompts efficiently using the queue:
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI( base_url="https://llm-server.llmhub.t-systems.net/queue",)
prompts = [ "Summarize the benefits of solar energy.", "Explain how wind turbines generate electricity.", "Describe the future of hydrogen fuel cells.",]
async def process_prompt(prompt): response = await client.chat.completions.create( model="gpt-oss-120b", messages=[{"role": "user", "content": prompt}], max_tokens=500, ) return response.choices[0].message.content
async def main(): tasks = [process_prompt(p) for p in prompts] results = await asyncio.gather(*tasks) for prompt, result in zip(prompts, results): print(f"Prompt: {prompt[:50]}...") print(f"Result: {result[:100]}...\n")
asyncio.run(main())import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://llm-server.llmhub.t-systems.net/queue",});
const prompts = [ "Summarize the benefits of solar energy.", "Explain how wind turbines generate electricity.", "Describe the future of hydrogen fuel cells.",];
const results = await Promise.all( prompts.map((prompt) => client.chat.completions.create({ model: "gpt-oss-120b", messages: [{ role: "user", content: prompt }], max_tokens: 500, }) ));
results.forEach((result, i) => { console.log(`Prompt: ${prompts[i]}`); console.log(`Result: ${result.choices[0].message.content}\n`);});Queue Endpoints for Other APIs
Section titled “Queue Endpoints for Other APIs”Embeddings
Section titled “Embeddings”from openai import OpenAI
client = OpenAI( base_url="https://llm-server.llmhub.t-systems.net/queue",)
response = client.embeddings.create( model="text-embedding-bge-m3", input="The benefits of renewable energy in Europe",)
print(f"Embedding dimension: {len(response.data[0].embedding)}")Audio Transcription
Section titled “Audio Transcription”from openai import OpenAI
client = OpenAI( base_url="https://llm-server.llmhub.t-systems.net/queue",)
with open("meeting_recording.mp3", "rb") as audio_file: transcript = client.audio.transcriptions.create( model="whisper-large-v3", file=audio_file, )
print(transcript.text)Image Generation
Section titled “Image Generation”from openai import OpenAI
client = OpenAI( base_url="https://llm-server.llmhub.t-systems.net/queue",)
result = client.images.generate( model="gpt-image-2", prompt="A futuristic data center powered by renewable energy",)
print(result.data[0].b64_json[:50] + "...")Next Steps
Section titled “Next Steps”- Chat Completions — Standard synchronous chat API
- Streaming — Real-time token-by-token responses
- Rate Limits — Understand TPM/RPM limits and headers
- API Endpoints — Full endpoint reference