Evidence vs. Opinion II
When the Numbers Agree
The two valuations landed close. That is when the real audit starts.
Two AI tools valued the same European bioethanol producer and landed within 8% of each other: $23.41 vs $25.33 per share. At first glance, a tie.
It was not. The number was never the product. The evidence underneath it was.
Claude for Excel
$23.41/share
Sharpe AI
$25.33/share
Market price
$12.18/share
Difference between tools
~8%
Sharpe record
205 pinned cells
Forecast ledger
16 drivers · 7 SAFE · 9 REVIEW
Claude for Excel
$23.41
One-tab model
Typed assumptions
Self-review
Chat caveats
Sharpe AI
$25.33
Five linked sheets
Tested drivers
Review flags
Bridge scan
Pinned workbook record
Same neighborhood. Different foundations.
The twist: the tools agreed.
The first case study was dramatic because the answers diverged.
This one is more dangerous because they converged.
Claude and Sharpe AI both looked at the same workbook and both concluded that the company was worth roughly twice its market price. Claude landed at $23.41/share. Sharpe AI landed at $25.33/share.
That is close enough that a casual reader might say:
Great. The tools agree. We are done.
That is exactly the wrong conclusion.
In finance, agreement is comforting. It is also dangerous. When two models agree, people stop asking questions.
This case study asks the question people usually skip:
What is underneath the agreement?
Agreement is nice. Blind agreement is how very polished mistakes get promoted to final versions.
The headline numbers looked similar.
| Tool | Deliverable | WACC | Terminal method | Value/share | vs. market |
|---|---|---|---|---|---|
| Claude for Excel | 1-tab model, self-reviewed | 9.4% | Perpetuity growth at 2.0% | $23.41 | +92% |
| Sharpe AI | 5 linked sheets, pinned artifact | 4.82% | Exit 6.20x EV/EBITDA | $25.33 | +108% |
Both tools saw the same broad opportunity: the market price looked meaningfully below the DCF result.
Both also flagged the same tension. The implied valuation multiple sat above the company's own trading history. Claude raised this as a caveat. Sharpe AI printed the warning into the model itself.
That difference matters.
A caveat in chat is useful.
A caveat inside the workbook is reviewable.
A similar answer is not the same thing as a similar model.
A similar answer can hide very different logic.
Two models can produce similar final values for completely different reasons.
One model may use a higher discount rate and a generous terminal growth assumption. Another may use a lower evidence-sourced WACC and a more conservative exit multiple. The final valuation can land in the same range even though the logic underneath is different.
That is what happened here.
Claude used a 9.4% WACC and 2.0% perpetuity growth.
Sharpe AI used a 4.82% WACC and a 6.20x EV/EBITDA exit multiple sourced from the company's own multiple history.
The answer was similar.
The route was not.
Claude path
Higher WACC: 9.4%
Perpetuity growth: 2.0%
$23.41/share
Sharpe path
Evidence-built WACC: 4.82%
Market-history exit multiple: 6.20x
$25.33/share
Claude produced a reasonable first draft.
Claude did not produce nonsense.
It created a usable one-tab valuation, applied a conventional WACC, used a perpetuity-growth terminal value, created a sensitivity range, and raised a sensible warning that the implied exit multiple looked high relative to the company's own trading history.
That is a useful first draft.
But it was still a first draft.
The model relied on typed assumptions: revenue growth, margin fade, tax rate, capex as a percentage of revenue, working capital as a percentage of incremental revenue, and terminal growth.
Those assumptions may be reasonable. The issue is that the workbook itself did not prove them.
Claude's model logic
| Area | Claude treatment | Comment |
|---|---|---|
| Revenue | Typed growth fade | Reasonable, but not tested line by line |
| Margin | Typed EBITDA margin fade | Sensible judgment, but not evidence-scored |
| Capex | Percentage of revenue | Clean assumption, but not validated against the workbook |
| Working capital | Percentage of incremental revenue | Simple, but hides line-item behavior |
| Terminal value | 2.0% perpetuity growth | Conventional, but material to value |
| Review | Self-reviewed | Useful, but not a pinned workbook record |
A good first draft is still a first draft. It does not become a reviewed model because it speaks confidently.
Sharpe AI produced a record, not just a result.
Sharpe AI also produced a valuation.
But the important part was the record behind it.
The run produced five linked sheets: DCF output, PP&E schedule, working capital schedule, cost of capital, and sensitivity analysis. The final artifact was pinned to the compiled workbook: 205 cells were substituted from the sheet's own evaluated values, so the chat, artifact, and Excel output could not quietly disagree.
That is the core product idea.
Not:
Here is a number.
But:
Here is the number, here is the workbook logic behind it, here is what was tested, here is what was flagged, and here is what still deserves review.
Workbook
Driver tests
TestedLinked DCF sheets
LinkedBridge scan
ScannedReconciliation checks
VerifiedPinned record
PinnedThe forecast ledger was the product.
Sharpe AI committed 16 forecast drivers. Each line carried a method, a backtest score, and a risk label.
Seven lines passed as SAFE
Nine lines were marked REVIEW
That is not a weakness. That is the point.
A serious model should not pretend every line is equally knowable. Some lines are stable. Some lines are noisy. Some lines need management guidance. Sharpe AI makes that visible.
How Sharpe AI tested a forecast method
Forecast test = method error ÷ baseline error
Before a method is used, Sharpe AI asks: would this method have worked on the company's own past data?
It compares the method's error against a simple baseline, such as repeating last year's number.
- Below 1 means the method beat the simple baseline.
- Above 1 means the method did not beat the baseline.
- A line can still be used if it is the best available method, but it is marked for review.
Forecast evidence examples
| Line item | Sharpe method | Score | Decision | Why it matters |
|---|---|---|---|---|
| COGS | Ratio to revenue | 0.192 | SAFE | Strong evidence that cost structure tracked revenue |
| Ethanol segment revenue | Growth path | 0.595 | SAFE | Segment growth had historical support |
| Bioethanol / food / feed revenue | Growth path | 2.286 | REVIEW | The line was not reliably forecastable from history |
| Depreciation & amortization | Growth path | 0.855 | REVIEW | Missed the bar by 0.005, so it was not rounded into a pass |
| Capex | Growth path | 1.153 | REVIEW | Lumpy line; needs human review or management input |
A weak number is allowed through. A hidden one is not.
The record mattered more than the point estimate.
Sharpe AI's final answer was not just a chat response.
It was a pinned artifact.
That means the final reported values came from the compiled workbook itself, not from a separate narrative summary. The record contained 51 lines, 17 assumptions, 15 reconciliation checks, and 205 substituted cells from the sheet's evaluated values.
This is critical.
In financial modelling, the chat should not be able to say one thing while the spreadsheet says another.
Chat answer
Generated artifact
Compiled workbook
205 cells substituted from the sheet's own evaluated values
The workbook and the chat were no longer allowed to have different opinions. Very rude to the model. Very useful to the analyst.
A pension deficit is invisible until something makes silence illegal.
The equity bridge is where small omissions become real valuation problems.
Claude's bridge used net cash.
Sharpe AI's bridge scanned the balance sheet and included net cash, associates, and the pension deficit.
The pension deficit was not huge in this case, but the lesson is important. The item exists whether the model looks for it or not.
A model that does not scan for claims can still be arithmetically clean.
It can also be incomplete.
Equity bridge comparison
| Bridge item | Claude | Sharpe AI | Why it matters |
|---|---|---|---|
| Net cash | Included | Included | Basic bridge item |
| Associates | Not separately surfaced | Included | Adds value if economically separate |
| Pension deficit | Not surfaced | Subtracted | Claim on equity value |
| Bridge coverage | Net debt only | Full claim scan | Prevents silent omissions |
Bridge completeness is not a style preference.
What went right. What still needed judgment.
What went right
- Claude produced a reasonable first-pass valuation.
- Both tools flagged the broad valuation tension versus market trading history.
- Sharpe AI generated a full linked model rather than a single narrative answer.
- Sharpe AI tested forecast lines and labelled weak ones instead of hiding them.
- The final Sharpe artifact was pinned to the workbook.
What still needed judgment
- Sharpe's 4.82% WACC is evidence-sourced but likely to be challenged by a committee.
- Claude's 9.4% WACC looks more conventional but is less traceable.
- Capex and tax were not cleanly forecastable from history.
- The stock trading far below both valuation outputs raises a real market question: what does the market know?
Evidence does not remove judgment. It tells you where judgment is still required.
The number was never the product.
This case is powerful because the final valuations were close.
If Sharpe AI only mattered when Claude was far away, the product would be weaker.
But that is not the point.
The value of Sharpe AI is not simply "a different answer."
The value is a controlled modelling workflow: evidence-backed assumptions, tested forecast drivers, linked schedules, bridge completeness, review flags, and a final record that matches the workbook.
When the numbers disagree, Sharpe AI helps explain why.
When the numbers agree, Sharpe AI helps prove what the agreement is worth.
You are not buying the answer. You are buying the ability to show your work.
Final takeaway
Agreement is not enough.
In finance, a model does not become decision-ready because two tools produced similar outputs. It becomes decision-ready when the assumptions can be traced, the mechanics can be checked, the bridge can be reconciled, and the remaining judgment calls are visible.
That is what Sharpe AI is being built to do.
Want to see what Sharpe AI would do with your model?
Bring a DCF, a messy workbook, a first draft from AI, or a model your team needs to defend.
For product demonstration only. Not investment advice.