lawless-m avatar

qwen-ollama

Using Qwen 2.5 models via Ollama for local LLM inference, text analysis, and AI-powered automation

作者 lawless-m|オープンソース

Qwen via Ollama

  • Chosen model: qwen2.5:7b (4.7 GB). Other sizes exist if needed: 0.5b / 14b / 32b.
  • Endpoint: http://localhost:11434 (standard Ollama API).
  • Timeouts: 120 s default for analysis calls; longer (e.g. 300 s) for heavy generation.
  • Prefer local over cloud when: offline, data can't leave the machine, or bulk work where API cost matters.
  • VRAM contention with other GPU services: see the Vram-GPU-OOM-memory-management skill.

Reference implementation (Marvinous)

  • OllamaClient.rs in this folder — copy of the production async client (reqwest, timeout, error handling) from /home/matt/Marvinous/src/llm/client.rs.
  • Prompt building: /home/matt/Marvinous/src/llm/prompt.rs
  • System prompt (Marvin's personality): /etc/marvinous/system-prompt.txt
  • Config (model/endpoint settings): /etc/marvinous/marvinous.toml
qwen-ollama - Claude Code・Cursor 対応の AIエージェント Skill | Agent Skills