yera.models.interfaces.llms.base

Base interface for LLM implementations.

This module defines the abstract BaseLLMInterface that all llm provider implementations must inherit from. It establishes the contract for:

  • Streaming chat completions (chat method)
  • Generating structured outputs conforming to a schema (make_struct method)
  • Managing llm client lifecycle (start/stop methods)

Concrete implementations (e.g., AnthropicLLM, OpenAILLM, AwsBedrockLLM) provide provider-specific implementations of these abstract methods whilst handling their respective API clients and configuration.

Symbols

class BaseLLMInterface — Abstract base interface for llm implementations.
class LLMToken — Class representing a token in an LLM response.
class RateLimitError — A provider rate-limit response that may succeed after waiting.

BaseLLMInterface

Inherits: ABC

Abstract base interface for llm implementations.

Defines the contract that all concrete llm provider implementations must satisfy. Subclasses handle provider-specific client initialisation, authentication, and API interaction whilst conforming to the streaming chat and structured output methods defined here.

Methods

chat — Stream a chat completion response.
make_struct — Stream a structured output response conforming to a provided schema.
make_request_struct — Generate a tool-like request structure with an initial `call_id` token.
start — Initialise the llm client.
stop — Shut down and clear the llm client.
with_instruction — Prepend or insert an instruction into the conversation history.

BaseLLMInterface.chat

chat(
    messages: list[Message],
    reasoning_level: ReasoningLevel | None = None,
    **kwargs,
) → Iterator[LLMToken]

Stream a chat completion response.

Abstract method that must be implemented by concrete llm providers. Sends a conversation to the llm and streams the response as text tokens.

Parameters

messages
type: list[Message]

List of Message objects representing the conversation history.

reasoning_level
type: ReasoningLevel | None = None

Set the reasoning effort level overriding the default (medium)

**kwargs
type: str | int | float | bool

Provider-specific keyword arguments to customise llm behaviour (e.g., temperature, top_p, max_tokens).

BaseLLMInterface.make_struct

make_struct(
    messages: list[Message],
    reasoning_level: ReasoningLevel | None = None,
    **kwargs,
) → Iterator[LLMToken]

Stream a structured output response conforming to a provided schema.

Abstract method that must be implemented by concrete llm providers. Generates a response that strictly conforms to the structure defined by the provided schema class.

Parameters

messages
type: list[Message]

List of Message objects representing the conversation history.

cls
type: type[TStruct]

A pydantic model class defining the output structure. The provider will transform this into the format required by its respective API.

reasoning_level
type: ReasoningLevel | None = None

Set the reasoning effort level overriding the default (off)

**kwargs
type: str | float | int | bool

Provider-specific keyword arguments to customise llm behaviour.

BaseLLMInterface.make_request_struct

make_request_struct(
    messages: list[Message],
    reasoning_level: ReasoningLevel | None = None,
    **kwargs,
) → Iterator[LLMToken]

Generate a tool-like request structure with an initial call_id token.

Prepares a unique identifier for the tool call and streams structured output tokens. This method is intended for tools that require a tool_call → tool_result interaction pattern, where the call_id must be sent first to associate results.

Parameters

messages
type: list[Message]

Conversation history.

cls
type: type[TStruct]

Struct subclass defining the tool's input/output schema.

reasoning_level
type: ReasoningLevel | None = None

Set the reasoning effort level overriding the default (off)

**kwargs
type: str | float | int | bool

Per-call LLM overrides (e.g., temperature, max_tokens).

BaseLLMInterface.start

start() → None

Initialise the llm client.

Lifecycle method called before making API requests. Concrete implementations override this to instantiate and configure their provider-specific client. Default implementation does nothing.

BaseLLMInterface.stop

stop() → None

Shut down and clear the llm client.

Lifecycle method called when finished with the llm. Concrete implementations override this to release resources and clean up the provider-specific client. Default implementation does nothing.

BaseLLMInterface.with_instruction

with_instruction(
    messages: list[Message],
    instruction: str | None,
) → list[Message]

Prepend or insert an instruction into the conversation history.

Parameters

messages
type: list[Message]

The existing conversation history.

instruction
type: str | None

The extra instruction to inject.

Returns

type: list[Message]

A new list of messages with the instruction incorporated.

LLMToken

Class representing a token in an LLM response.

Can be either "thinking" from a reasoning (CoT) trace, or "response" as in a normal response to the user.

Attributes

kind
type: Literal['response', 'thinking', 'call_id']

a string that defines this as a thinking or response token.

content
type: str

the content of the token.

RateLimitError

Inherits: RuntimeError

A provider rate-limit response that may succeed after waiting.