LazuliQ 1.5 Photon

A 4B-parameter model on a newer base, fine-tuned on 13,362 verified examples to reason more steadily and call tools. It still won't match the large assistants, and that isn't the point. Small models that run anywhere are where a lot of this is heading.

Early beta. Photon is an experiment and will get things wrong.

13,362verified training examples
7.8xthe data behind 1.2 Photon
Multimodalreads text, images and video

What changed in 1.5

Photon 1.5 is built on Qwen3.5-4B instead of Qwen2.5-3B-Instruct, and it learns from a dataset about eight times larger. The base model's ceiling is still the ceiling, but the fine-tune now has far more to work with.

Tools it knows when to use

Trained on 495 rows that offer a calculator, a clock and a unit converter, holding 395 real tool calls, plus examples where the right move is to answer directly. 1.2 never saw a tool call.

Reasoning-first data

77.85% of training rows are think-mode, up from 19.5% in 1.2, and answers are checked by executable verifiers before they enter the dataset.

Measured, not just trained

1.5 holds out validation and test data and keeps the best checkpoint. 1.2 trained on everything and was spot-checked with three prompts.

1.2 Photon vs 1.5 Photon

The differences that matter most, taken from each model's training code and dataset.

1.5 Photon
1.2 Photon
Model
ReleaseMonth of launch
October 20267 months after 1.2
March 2026
Base modelStarting point
Qwen3.5-4B
Qwen2.5-3B-Instruct
ParametersLanguage model
4B
3.09B
Native contextTokens
262,1448x longer
32,768
Multimodal visionImage understanding
Text, images and videoVision encoder kept at full precision
Text only
Benchmarks
ARC-Challenge25-shot accuracy
67.66%+9.38 points
58.28%
IFEvalStrict, zero-shot, 541 prompts
68.39%+22.73 points
45.66%
Training
Fine-tuningMethod
16-bit LoRA
QLoRA, 4-bit base
Training examplesDataset size
13,3627.8x more
1,711
Held-out dataValidation and test
3,1181,694 validation, 1,424 test
None
Answer checkingHow examples are made
Executable verifiers14,785 candidates checked
Hand-written pairs
Data
Think-mode rowsShare of training data
77.85%
19.5%
Tool-use examplesTraining rows with tools
495Calculator, clock, units
0
Multi-turn examplesTraining rows, 2+ user turns
477
0
Release
Quantized releaseGGUF for local chat
Q4_K_M GGUF4-bit, with an F16 vision projector
Q5_K_M GGUF5-bit

Bold marks the stronger value where a comparison is meaningful. ARC-Challenge is 25-shot. IFEval is zero-shot with strict grading on 541 prompts, and the two models use their own chat templates. The Q4_K_M GGUF is the release format, and it is verified to run text, image and video inference. 1.2 figures come from running its dataset script. 1.5 figures come from the 1.5 training repository, and data mix counts use its shipped reference bundle.

The training data, in numbers

Two benchmarks, then four measurements of what each model learned from.

ARC-Challenge

25-shot accuracy
1.5 Photon1.2 Photon
67.66%
58.28%
0%20%40%60%
Accuracy

IFEval

Strict instruction following, zero-shot
1.5 Photon1.2 Photon
68.39%
45.66%
0%20%40%60%
Accuracy on 541 prompts

Training examples

Dataset size
1.5 Photon1.2 Photon
13,362
1,711
05,00010,000
Examples

Think-mode rows

Share of training data
1.5 Photon1.2 Photon
77.85%
19.5%
025%50%75%
% of training rows

Coding and data tasks

Share of training data
1.5 Photon1.2 Photon
38.9%
0.9%
010%20%30%
% of training rows

Tool-use examples

Training rows with tools
1.5 Photon1.2 Photon
495
0
0100200300400
Rows

ARC-Challenge: 25-shot. IFEval: 541 prompts, zero-shot, strict grading, each model with its own chat template. Dataset figures: 1.2 figures come from running its dataset script (1,711 rows). 1.5 totals are from the v6.9 build (13,362 rows); share figures use the shipped reference bundle (10,106 training rows).

A custom agentic harness

LazuliQ doesn't just chat. It runs inside a full custom agentic harness that turns a request into a visible plan, then does the work: it edits files, creates tasks, checks its own output and builds real files you can download. The model writes the content; the harness decides what happens next.

Edits files

Websites, PDFs, Word files, slides, Markdown, JSON, CSV, CSS, JS and Python. Follow-ups change the latest version, and each edit is kept as its own revision.

Creates tasks

Every request becomes a short plan shown as a timeline, with the line the model is writing right now and what each check found.

Checks its own work

Loops, placeholders, topic drift and broken files are caught by deterministic checks. Small fixes are made locally; the rest go back to the model with the exact complaint.

Uses tools

An exact calculator, Wikipedia search with citations, file reading and workspace commands. The harness does the deciding, so the small model only has to write.

Still small and often wrong, so check anything that matters. See it work in the demo.

StartumRAG

A retrieval layer that runs server-side and offline, giving a small model access to facts it was never big enough to hold. No external tool calls, no waiting on the network. It's the most experimental part of Photon, and the most interesting.

Offline and private

Runs inside isolated, internet-free server environments, so nothing leaves the machine. On-device deployment is next.

Ultra-low latency

With no external web APIs in the loop, retrieval adds very little time on top of inference.

Fewer invented answers

Grounding answers in real sources cuts down the guessing a small model would otherwise do when it hits the edge of what it knows.

Under the hood

A hybrid decoder inherited from Qwen3.5-4B, mixing linear attention with full attention for efficiency at long context.

Total parameters
4B
Context length
262K
Layers
32
Vocabulary
248K

Gated DeltaNet and full attention

24 of 32 layers use linear attention and every fourth layer uses full gated attention, with 16 query and 4 key-value heads.

16-bit LoRA on 200 projections

Rank 16 adapters cover 72 linear-attention, 32 full-attention and 96 MLP projections, verified before training starts.

A multimodal Apache-2.0 base

Qwen3.5-4B includes a vision encoder. The fine-tune trains the text layers only, and the merge check confirms the vision weights are unchanged.

Ships as Q4_K_M GGUF

Language weights run in 4-bit with an F16 vision projector, so text, images and video work on CPU. The export is verified to run text, image and video before it is published.

About this release

LazuliQ 1.5 Photon is a student project: a fine-tune of an open-source 4B base model, built and maintained by one developer. It's an experiment in how capable a small model can be, not a replacement for the large assistants you already use. This page compares training recipes and data, and reports two benchmarks, ARC-Challenge and IFEval; it makes no wider quality claims. Expect mistakes: the model can be wrong, invent details, or lose the thread of a long conversation. Everything here is beta, StartumRAG most of all, so check anything important before relying on it.

Built by one student

DominikWalser

AI & Technology Enthusiast

Startups & Innovation

LazuliQ is an independent project, designed, researched and developed by a single student developer with a passion for AI, LLMs and cybersecurity. Open to collaboration, research partnerships and new opportunities.

Artificial Intelligence LLMs Machine Learning Mobile Cybersecurity Startups & Innovation

See what 1.5 can do