Slim::Options
Sampling options for text generation.
These are standard LLM parameters supported by most inference engines. They control how the model selects the next token during generation.
Common Parameters
temperature- Controls randomness. 0 = deterministic, 1 = creativetop_p- Nucleus sampling. Consider tokens with cumulative probability ptop_k- Only consider the k most likely tokensrepeat_penalty- Penalize repeated tokens (reduces repetition)seed- Random seed for reproducible outputs
Ollama-Specific Parameters
These are passed through but are specific to the Ollama/llama.cpp backend:
num_ctx- Context window size (default: 2048)num_predict- Max tokens to generate (default: 128, -1 = infinite)stop- Stop sequences to end generation
Example
# For classification (deterministic)
opts = Slim::Options.new(temperature: 0.1, top_k: 1)
# For creative generation
opts = Slim::Options.new(temperature: 0.8, top_p: 0.9)
response = client.generate("llama3.2:3b", prompt, options: opts)
Constructors
Instance methods
Maximum tokens to generate. -1 for unlimited. Default: 128
Penalize tokens that have appeared. Higher = less repetition. Range: 0.0-2.0. Default: 1.1
Penalize tokens that have appeared. Higher = less repetition. Range: 0.0-2.0. Default: 1.1
Stop sequences - generation stops when these are encountered.
Randomness of output. 0.0 = deterministic, 1.0 = creative, 2.0 = chaotic. For classification tasks, use 0.0-0.3. Default: 0.8
Randomness of output. 0.0 = deterministic, 1.0 = creative, 2.0 = chaotic. For classification tasks, use 0.0-0.3. Default: 0.8
Nucleus sampling threshold. Only consider tokens whose cumulative probability mass reaches this value. Range: 0.0-1.0. Default: 0.9