On August 13, 2026, Investor D stopped running on Gemini 3.1 Pro and started running on Gemini 3.7 Flash.
One of the four seats on our Investment Committee changed the engine it reasons with. That is a bigger event than it sounds, because a seat is not a person who learned something. It is a model, and swapping it means every verdict after the swap comes from a different reasoner than every verdict before it.
So the change is marked in the log. Anyone scoring this committee later can see exactly where the seam is and refuse to average across it.
What changed
Investor D is the growth and runway seat, the one that classifies a business by what kind of grower it is and asks how much category is left. It reads the same fact sheet as the other three and cannot see their work.
It now runs on Gemini 3.7 Flash at high reasoning effort. It previously ran on Gemini 3.1 Pro.
The version number goes up and the tier goes down. That combination is the reason we tested it instead of taking the upgrade on faith.
Why we did not simply take the newer model
We keep a rule that anything reaching the public site is not a flash tier task. It was written after a flash tier model failed the same publishing job twice in a row, emitting malformed syntax and exiting cleanly having written nothing. A committee seat's reasoning goes straight into published articles, so on the face of it that rule should have ended this conversation.
It does not apply here, and the reason is worth stating precisely rather than waving through.
That failure was a tool calling failure. The model was asked to drive a sequence of actions and produced action syntax that could not be executed. An investor seat has no tools at all. Its instructions say so in as many words: no tools, no helper scripts, reason only from the fact sheet in front of you. There is no tool call for a flash tier model to malform, because there are no tool calls.
The rule still holds everywhere else, unchanged.
The test
One run on the old model against three runs on the new one, on an identical prompt: the seat's philosophy verbatim, one NVIDIA fact sheet built the same morning, and the same deliverable instructions. Nothing else differed.
| Gemini 3.1 Pro | Gemini 3.7 Flash | |
|---|---|---|
| Verdict | Hold, conviction 7 | Hold, conviction 7, all three runs |
| Time to answer | 73 seconds | 23, 25 and 28 seconds |
| Metrics cited from the sheet | 5 | 15 |
| Figures not traceable to the sheet | none | none |
Four runs, one verdict, one conviction score. The new model is roughly three times faster and reads considerably more of the sheet, citing fifteen figures where the old one cited five.
One figure in the new model's answer looked wrong and was not. It reported total debt of $8.47B, which appears nowhere on the sheet. The sheet carries long term debt of $7.47B and short term debt of $1.00B, and the model showed both components in the same sentence. That is addition, not invention, and the distinction is the whole job when you are checking a machine's arithmetic.
What the test did not show
We went in expecting to find that the cheaper model was sloppier, and we did not find it.
Both models round entry levels to numbers that are not on the fact sheet. The new one did it in all three runs. The old one did it in its August 10 committee answer, which named a range no line of that sheet supports. Neither is better behaved than the other here.
Both models also sometimes write their verdict heading with a hyphen before the word, which until yesterday caused our log to record the verdict as blank. The new model did it once in three runs. The old model did it in three of its nine logged sessions. That defect belonged to the log rather than to either model, and we wrote about finding it the same day it was fixed. The parser now accepts any separator, reads three heading styles, and requires word boundaries, so a seat writing "no buyer at this price" can no longer be recorded as having voted Buy.
Four runs is not a study. It is enough to show that the substance did not move and not enough to rank the two models on judgment. We are not claiming the new seat is smarter. We are claiming it reached the same conclusion, faster, on the one case we could test properly.
The seam
The verdict log now carries a line that is not a verdict. It records which seat changed, which model it left, which model it joined, why, and one instruction: do not score across this boundary as one reasoner.
Ten sessions were logged before the change. Every Investor D verdict in them came from the old model. That fact is now impossible to lose, which was the entire point of writing it down at the moment it happened rather than reconstructing it later from memory.
There is a second effect worth recording. Our technical commentary layer, which runs after a committee decision and has no vote, sits on Gemini 3.1 Pro. Until this change it shared a model with Investor D, and that was tolerable only because it does not vote: two correlated voices inside a four seat count would quietly halve the committee's independence. After the change the overlap is gone and the technical layer is independent of every seat.
That is an improvement we did not plan and did not earn. We have written the argument into both files anyway, because if either one is ever moved the overlap comes back, and the reasoning that made it survivable needs to be sitting there when someone does it.
What this does not establish
It does not establish that Investor D is better. It establishes that on one company, on one day, across four runs, it gave the same answer in less time.
It does not establish that the seats disagree usefully. No verdict in the log has been scored against an outcome yet, on either model.
And it does not make the two halves of the record comparable. A seam that is documented is still a seam. When there is finally enough history to judge this committee, Investor D's will be in two pieces, and the honest way to read it is as two pieces.
As of August 13, 2026, BagsCapital manages $0 in assets under management, holds no position in any asset named here, and has never accepted outside capital. The committee is advisory and does not trade. Nothing in this note is a recommendation.