How vLLM requests work
vLLM workers are queue-based Serverless endpoints. They use the same/run and /runsync operations as other Runpod endpoints, following the standard Serverless request structure.
The key difference is the input format. vLLM workers expect specific parameters for language model inference, such as prompts, messages, and sampling parameters. The worker’s handler processes these inputs using the vLLM engine and returns generated text.
Request operations
vLLM endpoints support both synchronous and asynchronous requests.Asynchronous requests with /run
Use /run to submit a job that processes in the background. You’ll receive a job ID immediately, then poll for results using the /status endpoint.
Synchronous requests with /runsync
Use /runsync to wait for the complete response in a single request. The client blocks until processing is complete.
Input formats
vLLM workers accept two input formats for text generation.Messages format (for chat models)
Use the messages format for instruction-tuned models that expect conversation history. The worker automatically applies the model’s chat template.Prompt format (for text completion)
Use the prompt format for base models or when you want to provide raw text without a chat template.Applying chat templates to prompts
If you use the prompt format but want the model’s chat template applied, setapply_chat_template to true.
Request input parameters
Here are all available parameters you can include in theinput object of your request.
Sampling parameters
Sampling parameters control how the model generates text. Include them in thesampling_params dictionary in your request.