Use GenericChatCompletionRequest with GenericChatCompletionLanguageModel to send text conversations to language models through a common API. The model input selects the model, and the request contains the messages and generation parameters.
This page documents the Python classes from palantir_models and language-model-service-api. For repository setup and model selection, see Use language models within transforms.
Add palantir_models to your repository. Replace <model-rid> with the full resource identifier of a model you can access, and replace the output dataset path. This example uses a lightweight transform.
Copied!1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31import pandas as pd from language_model_service_api.languagemodelservice_api_completion_v3 import ( DialogChatMessage, DialogRole, GenericChatCompletionRequest, ) from palantir_models.models import GenericChatCompletionLanguageModel from palantir_models.transforms import GenericChatCompletionLanguageModelInput from transforms.api import Output, transform @transform.using( model=GenericChatCompletionLanguageModelInput("<model-rid>"), output=Output("/path/to/output/dataset"), ) def compute(model: GenericChatCompletionLanguageModel, output): request = GenericChatCompletionRequest( messages=[ DialogChatMessage( role=DialogRole.SYSTEM, content="Answer questions in one sentence.", ), DialogChatMessage( role=DialogRole.USER, content="Why is the sky blue?", ), ], max_tokens=256, ) response = model.create_chat_completion(request) output.write_table(pd.DataFrame({"completion": [response.completion]}))
GenericChatCompletionLanguageModelInput supplies a GenericChatCompletionLanguageModel to the compute function. The same model input and request types can be used in a Spark transform with @transform(...).
GenericChatCompletionRequestImport GenericChatCompletionRequest from language_model_service_api.languagemodelservice_api_completion_v3.
| Python parameter | Type | Required | Description |
|---|---|---|---|
messages | list[DialogChatMessage] | Yes | The conversation to send to the model, in order. See Message structure. |
temperature | float | No | Controls sampling randomness where supported. Supported values and behavior depend on the selected model. |
stop_sequences | list[str] | No | Sequences that stop generation when encountered, where supported by the model. |
max_tokens | int | No | Maximum number of output tokens to generate for this request. This excludes input tokens and is subject to the selected model's limits. |
The optional parameters can be omitted. There is no single set of defaults or supported parameter values across all models; the service and selected model determine behavior for omitted parameters. Start with the required messages and add only parameters supported by your model.
The Python names are stop_sequences and max_tokens. In the serialized request, these fields are named stopSequences and maxTokens.
This request supports text messages and the generation parameters listed above. It does not define image inputs, tool calls, or a structured output schema.
Import DialogChatMessage and DialogRole from the same module as GenericChatCompletionRequest.
| Field | Type | Description |
|---|---|---|
role | DialogRole | DialogRole.SYSTEM, DialogRole.USER, or DialogRole.ASSISTANT. |
content | str | The text of the message. |
Follow these rules when constructing a conversation:
SYSTEM or USER message.SYSTEM only as the first message, if needed.SYSTEM message with a USER message.USER and ASSISTANT messages after the optional system message.For example, a conversation with history can have the roles SYSTEM, USER, ASSISTANT, USER. Include the previous turns in messages when you want the model to use them as context.
GenericChatCompletionResponsemodel.create_chat_completion(request) returns a GenericChatCompletionResponse with these fields:
| Python attribute | Type | Description |
|---|---|---|
completion | str | The generated text. Read this field to retrieve the answer. |
token_usage | TokenUsage | Token counts and the model's context-window limit. |
The token usage object exposes prompt_tokens for input tokens, completion_tokens for output tokens, and max_tokens for the model's context window, which includes both input and output tokens. response.token_usage.max_tokens has a different meaning from the request's max_tokens output limit.
This response has a completion field rather than the choices list used by GptChatCompletionResponse.
Set the model RID on GenericChatCompletionLanguageModelInput. GenericChatCompletionRequest does not have a model or provider field.
For a registered model, use its full ri.language-model-service..registered-model.<id> RID. The Python client uses that identifier to call the registered model. Registering and configuring the model determines which model endpoint handles the request.
When investigating a custom model integration:
response.completion.For guidance on connecting a self-hosted or external model to AIP, see Bring your own model to AIP.