LazuliQ 1.5 Photon
A 4B-parameter model on a newer base, fine-tuned on 13,362 verified examples to reason more steadily and call tools. It still won't match the large assistants, and that isn't the point. Small models that run anywhere are where a lot of this is heading.
Early beta. Photon is an experiment and will get things wrong.
What changed in 1.5
Photon 1.5 is built on Qwen3.5-4B instead of Qwen2.5-3B-Instruct, and it learns from a dataset about eight times larger. The base model's ceiling is still the ceiling, but the fine-tune now has far more to work with.
Tools it knows when to use
Trained on 495 rows that offer a calculator, a clock and a unit converter, holding 395 real tool calls, plus examples where the right move is to answer directly. 1.2 never saw a tool call.
Reasoning-first data
77.85% of training rows are think-mode, up from 19.5% in 1.2, and answers are checked by executable verifiers before they enter the dataset.
Measured, not just trained
1.5 holds out validation and test data and keeps the best checkpoint. 1.2 trained on everything and was spot-checked with three prompts.
1.2 Photon vs 1.5 Photon
The differences that matter most, taken from each model's training code and dataset.
Bold marks the stronger value where a comparison is meaningful. The Q4_K_M GGUF is the release format, and it is verified to run text, image and video inference. Its answer quality is not benchmarked here. 1.2 figures come from running its dataset script. 1.5 figures come from the 1.5 training repository, and data mix counts use its shipped reference bundle.
The training data, in numbers
Four measurements of what each model learned from. These describe the data, not benchmark results.
Training examples
Think-mode rows
Coding and data tasks
Tool-use examples
1.2 figures come from running its dataset script (1,711 rows). 1.5 totals are from the v6.9 build (13,362 rows); share figures use the shipped reference bundle (10,106 training rows).
A custom agentic harness
LazuliQ doesn't just chat. It runs inside a full custom agentic harness that turns a request into a visible plan, then does the work: it edits files, creates tasks, checks its own output and builds real files you can download. The model writes the content; the harness decides what happens next.
Hi, I’m LazuliQ.
A small research model with an agent harness: I plan, write, check my own work and build real files. Still small and often wrong, so check anything that matters.
- !Plan the site
- !Write the content
- !Check the draft
- !Design and build
- !Final check
- !Change the palette
Slow coffee, made properly.
Small-batch roasts and fresh pastries at Harbour Street 12.
See the menuMenu
- Flat white€3.40
- Filter coffee€3.00
- Almond croissant€3.20
Edits files
Websites, PDFs, Word files, slides, Markdown, JSON, CSV, CSS, JS and Python. Follow-ups change the latest version, and each edit is kept as its own revision.
Creates tasks
Every request becomes a short plan shown as a timeline, with the line the model is writing right now and what each check found.
Checks its own work
Loops, placeholders, topic drift and broken files are caught by deterministic checks. Small fixes are made locally; the rest go back to the model with the exact complaint.
Uses tools
An exact calculator, Wikipedia search with citations, file reading and workspace commands. The harness does the deciding, so the small model only has to write.
Still small and often wrong, so check anything that matters. See it work in the demo.
StartumRAG
A retrieval layer that runs server-side and offline, giving a small model access to facts it was never big enough to hold. No external tool calls, no waiting on the network. It's the most experimental part of Photon, and the most interesting.
Offline and private
Runs inside isolated, internet-free server environments, so nothing leaves the machine. On-device deployment is next.
Ultra-low latency
With no external web APIs in the loop, retrieval adds very little time on top of inference.
Fewer invented answers
Grounding answers in real sources cuts down the guessing a small model would otherwise do when it hits the edge of what it knows.
Under the hood
A hybrid decoder inherited from Qwen3.5-4B, mixing linear attention with full attention for efficiency at long context.
- Total parameters
- 4B
- Context length
- 262K
- Layers
- 32
- Vocabulary
- 248K
Gated DeltaNet and full attention
24 of 32 layers use linear attention and every fourth layer uses full gated attention, with 16 query and 4 key-value heads.
16-bit LoRA on 200 projections
Rank 16 adapters cover 72 linear-attention, 32 full-attention and 96 MLP projections, verified before training starts.
A multimodal Apache-2.0 base
Qwen3.5-4B includes a vision encoder. The fine-tune trains the text layers only, and the merge check confirms the vision weights are unchanged.
Ships as Q4_K_M GGUF
Language weights run in 4-bit with an F16 vision projector, so text, images and video work on CPU. The export is verified to run text, image and video before it is published.
About this release
LazuliQ 1.5 Photon is a student project: a fine-tune of an open-source 4B base model, built and maintained by one developer. It's an experiment in how capable a small model can be, not a replacement for the large assistants you already use. This page compares training recipes and data, and reports no benchmark scores. Expect mistakes: the model can be wrong, invent details, or lose the thread of a long conversation. Everything here is beta, StartumRAG most of all, so check anything important before relying on it.
DominikWalser
AI & Technology Enthusiast
Startups & Innovation
LazuliQ is an independent project, designed, researched and developed by a single student developer with a passion for AI, LLMs and cybersecurity. Open to collaboration, research partnerships and new opportunities.