Complete reference documentation for the HuiLink API.
All API requests require authentication via a Bearer token in the Authorization header.
Authorization: Bearer sk-xxxxxxxxxxxxxxxxAuthorizationBearer sk-xxxxxxxxContent-Typeapplication/json| Method | Endpoint | Description |
|---|---|---|
POST | /v1/chat/completions | Send a chat completion request to a supported model. |
POST | /v1/responses | OpenAI Responses API — stateful, tool-native generation on the same model routing. |
POST | /v1/messages | Anthropic Messages protocol — run Claude-format clients against gateway models. |
POST | /v1/embeddings | Create an embedding vector representing the input text. |
POST | /v1/images/generations | Generate images from text prompts |
POST | /v1/video/generations | Generate videos from text prompts |
POST | /v1/audio/speech | Text-to-speech — returns binary audio |
POST | /v1/audio/transcriptions | Transcribe audio to text (multipart) |
POST | /v1/rerank | Rerank documents by relevance to a query |
GET | /v1/models | List the models available to your account. |
Send a chat completion request to a supported model.
curl https://gateway.hkting.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-api-key" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"temperature": 0.7,
"max_tokens": 1024
}'modelModel ID to use (e.g. gpt-4o, deepseek-v4-flash)
messagesArray of message objects with role and content
temperatureSampling temperature 0-2 (default 1.0)
max_tokensMaximum tokens to generate (default 4096)
streamEnable streaming via SSE (default false)
top_pNucleus sampling parameter (default 1.0)
frequency_penaltyPenalize token frequency (-2.0 to 2.0, default 0)
presence_penaltyPenalize token presence (-2.0 to 2.0, default 0)
stopUp to 4 sequences where the API stops generating
toolsList of tool definitions for function calling
userEnd-user identifier for monitoring
idUnique request identifier
objectAlways chat.completion
createdUnix timestamp of creation
modelThe model used for completion
choicesList of completion choices
usage.prompt_tokensNumber of prompt tokens
usage.completion_tokensNumber of completion tokens
usage.total_tokensTotal tokens consumed
{
"id": "chatcmpl-9a8b7c6d5e4f",
"object": "chat.completion",
"created": 1734567890,
"model": "gpt-4o",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 9,
"total_tokens": 21
}
}Create an embedding vector representing the input text.
curl https://gateway.hkting.com/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-your-api-key" \
-d '{
"model": "text-embedding-ada-002",
"input": "The quick brown fox jumps over the lazy dog"
}'modelEmbedding model ID (e.g. text-embedding-ada-002)
inputInput text or array of tokens to embed
{
"object": "list",
"data": [
{
"object": "embedding",
"index": 0,
"embedding": [
0.0023064255,
-0.009327292,
... // 1536 dimensions
]
}
],
"model": "text-embedding-ada-002",
"usage": {
"prompt_tokens": 8,
"total_tokens": 8
}
}List the models available to your account.
curl https://gateway.hkting.com/v1/models \
-H "Authorization: Bearer sk-your-api-key"| Model | Provider | Context Length | Features |
|---|---|---|---|
gpt-4o | OpenAI | 128K | Vision, Function Calling |
gpt-4o-mini | OpenAI | 128K | Vision |
gpt-4-turbo | OpenAI | 128K | Vision, Function Calling |
deepseek-v4-flash | DeepSeek | 1M | Reasoning, Fast |
deepseek-v4-pro | DeepSeek | 1M | Reasoning, Pro |
qwen-plus | Qwen | 128K | General |
qwen-max | Qwen | 128K | Function Calling |
Generate images from text descriptions using our media generation providers (Jimeng/即梦, DALL-E, etc.).
curl https://gateway.hkting.com/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jimeng-image",
"prompt": "A serene mountain landscape at sunset",
"n": 1,
"size": "1024x1024"
}'modelImage model ID (e.g. jimeng-image, dall-e-3)
promptText description of the desired image
nNumber of images to generate (1-4, default 1)
sizeImage size (e.g. 1024x1024, 1920x1080, default 1024x1024)
styleImage style (e.g. realistic, cartoon, anime) — supported by Jimeng
negative_promptWhat to avoid in the generated image
seedRandom seed for reproducible results
createdUnix timestamp of creation
dataArray of generated image objects
data[].urlURL of the generated image (expires after 24h)
data[].revised_promptAI-revised version of the input prompt (if applicable)
{
"created": 1734567890,
"data": [
{
"url": "https://.../generated-image.png",
"revised_prompt": "A serene mountain landscape at sunset with golden light..."
}
]
}| Model | Provider | Type |
|---|---|---|
jimeng-image | Jimeng (即梦) | Text-to-Image |
dall-e-3 | OpenAI | Text-to-Image |
dall-e-2 | OpenAI | Text-to-Image |
Generate short videos from text descriptions.
curl https://gateway.hkting.com/v1/video/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jimeng-video",
"prompt": "A cat walking in the park",
"duration": 5
}'modelVideo model ID (e.g. jimeng-video)
promptText description of the desired video
durationVideo duration in seconds (5 or 10, default 5)
sizeAspect ratio (e.g. 16:9, 9:16, 1:1, default 16:9)
negative_promptWhat to avoid in the generated video
seedRandom seed for reproducible results
idUnique video generation request ID
statusGeneration status (e.g. processing, completed, failed)
outputURL of the generated video file
{
"id": "video-gen-xxx",
"status": "completed",
"output": "https://.../generated-video.mp4"
}| Model | Provider | Max Duration |
|---|---|---|
jimeng-video | Jimeng (即梦) | 10s |
OpenAI-compatible audio endpoints. Speech turns text into audio (binary response); transcriptions convert uploaded audio to text via multipart/form-data (max 50 MB). Requests route to the first capable channel and fail over automatically.
curl https://gateway.hkting.com/v1/audio/speech \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Hello! Welcome to our service.",
"voice": "alloy",
"response_format": "mp3"
}' \
--output speech.mp3modelTTS model ID (e.g. tts-1, tts-1-hd)
inputText to synthesize (max ~4096 chars)
voiceVoice ID (e.g. alloy, echo, fable, onyx, nova, shimmer)
response_formatmp3 | opus | aac | flac | wav | pcm (default mp3)
speedSpeech speed, 0.25–4.0 (default 1.0)
The raw audio bytes with the upstream Content-Type (e.g. audio/mpeg). Save the body directly to a file.
curl https://gateway.hkting.com/v1/audio/transcriptions \
-H "Authorization: Bearer YOUR_API_KEY" \
-F file="@speech.mp3" \
-F model="whisper-1" \
-F response_format="json"fileAudio file (mp3, wav, m4a, webm, ...; max 50 MB)
modelTranscription model ID (e.g. whisper-1)
languageISO-639-1 hint (e.g. en, zh)
promptOptional spelling/style guidance
response_formatjson | verbose_json | text | srt | vtt (default json)
{
"text": "Hello! Welcome to our service."
}Cohere/SiliconFlow-compatible rerank: score a list of documents against a query by relevance. Documents may be plain strings or objects — both pass through unchanged.
curl https://gateway.hkting.com/v1/rerank \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bge-reranker-v2-m3",
"query": "What is machine learning?",
"documents": [
"Machine learning is a subset of AI.",
"The weather is nice today.",
"Deep learning uses neural networks."
],
"top_n": 2
}'modelRerank model ID (e.g. bge-reranker-v2-m3)
queryThe search query
documentsDocuments to rank — strings or objects
top_nReturn only the top N results (default: all)
idUnique request ID (if provided by upstream)
modelModel that processed the request
resultsRanked results, sorted by relevance (descending)
results[].indexPosition of the document in the input array
results[].relevance_scoreRelevance score against the query
results[].documentEchoed document (if the upstream returns it)
{
"id": "rerank-xxx",
"model": "bge-reranker-v2-m3",
"results": [
{ "index": 0, "relevance_score": 0.93 },
{ "index": 2, "relevance_score": 0.87 }
]
}API requests are subject to rate limits based on your plan. Exceeding limits returns a 429 status code.
| Plan | RPM | TPM |
|---|---|---|
| Free | 20 | 40,000 |
| Pay-as-you-go | 500 | 200,000 |
| Enterprise | Custom | Custom |
Rate limits can be configured per API key by an admin. Contact your admin to adjust limits.
The API uses standard HTTP status codes to indicate the outcome of a request.
| Code | Meaning | Common Cause |
|---|---|---|
200 | OK | Request succeeded |
400 | Bad Request | Invalid request body or parameters |
401 | Unauthorized | Invalid or missing API key |
402 | Payment Required | Insufficient balance |
404 | Not Found | Model or endpoint not found |
422 | Unprocessable Entity | Request validation failed |
429 | Too Many Requests | Rate limit exceeded |
500 | Internal Server Error | Internal server error |
502 | Bad Gateway | Upstream provider returned an error |
503 | Service Unavailable | All upstream providers are unavailable |