Using built-in presets

The built-in presets ship tuned defaults for every popular open-weight model family. Pick one, pass it to AsposeLLMApi.Create, and the engine handles model download, binary deployment, sampler tuning, and chat template selection.

This page helps you pick the right preset for your scenario and highlights the minimal code to use it. For full catalog details, see Supported presets.

Minimal usage

using Aspose.LLM;
using Aspose.LLM.Abstractions.Parameters.Presets;

var preset = new Qwen25Preset();
using var api = AsposeLLMApi.Create(preset);

string reply = await api.SendMessageAsync("Hello!");

Every preset follows the same pattern: swap the class name to change the model.

Picker by task

Text chat

Goal Preset Notes
Balanced general assistant Qwen25Preset, Qwen3Preset, Llama31_8BPreset, or Mistral7Preset 7-8B, good at most tasks.
Smallest footprint Llama32Preset (3B) or Phi4Preset (mini) Run on modest hardware.
Smallest possible (CPU-only) SmallModelPreset (0.5B), TinyLlamaPreset (1.1B), or Llama32_1BPreset (1B) Tutorials, smoke tests, edge boxes.
Very long context Llama32Preset (131K) or DeepSeekCoder2Preset (163K) For long documents.
Coding DeepSeekCoder2Preset, Qwen25Coder7BPreset, Qwen3Coder30BPreset, or DevstralSmall2_24BPreset Specialized training on code.
Multilingual coverage Qwen3Preset or MistralSmall3Preset Trained on broad language mixes.
Enterprise-tuned Granite3_8BPreset IBM Granite 3.1, safety-aligned.
Fully-open research Olmo2_7BPreset AllenAI OLMo 2, fully open training and data.
Frontier on a workstation Llama3_3_70BPreset, SeedOss36BPreset, or MistralSmall3Preset 24-70B, needs 32 GB+ RAM and partial GPU offload.
Step-by-step reasoning DeepseekR1Qwen3Preset, Ernie4_5_21BPreset, or SeedOss36BPreset Chain-of-thought style output. Budget 1024-2048 MaxTokens.
No GPU available The *PresetCpu twin of any preset: Mistral7PresetCpu, Llama31_8BPresetCpu, Phi35MiniPresetCpu, … 27 twins ship. GpuLayers = 0, context capped at 4 K. See CPU-tuned variants.
No built-in preset: use your own GGUF Extend PresetCoreBase See Creating from scratch.

Vision

Goal Preset Notes
Smallest vision model with long context Qwen3VL2BPreset 2B, 262K context.
Strongest reasoning on images Ministral3VisionPreset 8B, 262K.
Vision + chain-of-thought reasoning NemotronOmniPreset 30B MoE, multimodal + reasoning; needs 32 GB+ RAM.
Lightweight GLM-family vision Glm4_6VFlashPreset Zhipu GLM-4.6V Flash.

See Vision presets for details and memory requirements.

Trade-offs

Dimension Smaller preset Larger preset
Speed Faster (more tokens/sec) Slower
Memory Less RAM/VRAM More
Quality Lower on complex tasks Higher
Cost (machine time) Lower Higher

Start with the smallest preset that meets your quality bar. Move up only if the output is not good enough.

Common overrides on built-in presets

Tweak defaults before Create without changing the preset class:

var preset = new Qwen25Preset();

// Make output more deterministic.
preset.SamplerParameters.Temperature = 0.2f;

// Use a smaller context to save memory.
preset.ContextParameters.ContextSize = 8192;

// Set a default system prompt.
preset.ChatParameters.SystemPrompt = "You are a concise assistant.";

using var api = AsposeLLMApi.Create(preset);

See Customizing for the full pattern and common knobs.

Before you ship

What’s next