The first version of Quant was, by any reasonable measure, a good writer and a bad analyst. Ask it whether a name was extended and it produced three fluent paragraphs of technical vocabulary that were not anchored to a single number on the chart in front of you. It was pattern-matching the genre of market commentary rather than reading the market.
The fix was not a bigger model. It was refusing to let the model see prose in the first place. Quant now receives a structured snapshot of exactly what your chart shows: the visible bar range, the series of every indicator you have applied with its parameters, the session boundaries, the volume profile, and the corporate events that fall inside the window. No narrative summary, no pre-digested "the stock is bullish". Just typed series with timestamps.
On top of that we run a small set of deterministic tools the model may call: fetch a window of bars, compute an indicator it does not already have, pull the last N earnings prints, or run a screener query. Every call is logged, and the response carries a citation token that ties each returned number to its source. When Quant states that a stock closed above its 50-day average on eleven of the last thirteen sessions, the interface can highlight those eleven bars, because the claim is welded to the rows that produced it.
Citation changed the failure mode entirely. A model that must attach a token to every quantitative claim cannot smoothly invent one, because there is no token to attach. In our internal evaluation set of 1,200 questions with verifiable answers, unsupported numeric claims fell from roughly 18% of responses to under 2%. The remaining failures are mostly arithmetic on multi-step comparisons, which we now hand to a tool rather than to the model.
We also had to teach it to refuse. Quant will not tell you what to buy, will not project a price target, and will say plainly when a question falls outside the data you are entitled to rather than improvising an answer from general knowledge. That refusal is a product decision as much as a safety one: an assistant that occasionally says "the data does not support an answer here" is one you can actually build a process around.
The last piece was latency. Chart context is large, and re-serialising it on every turn made follow-up questions slow and expensive. We now hash the context, cache the encoded form for the life of the thread and send only deltas as you pan or add an indicator. Median response time on a five-year daily series went from 4.1 seconds to 1.6, which is the difference between a tool you consult and a tool you converse with.
Dev Ranganathan
Principal Engineer, AI
Writes for AlgoBeam on engineering. Every figure quoted above can be reproduced in the backtester with the same settings.