Skip to content

Jose AI Assistant, tested.

What the testing showed

Testing caught an unfair leap in how the assistant described my fit, confirmed that the fix held across repeated runs, and showed that the cheaper model performed just as well.

I didn’t want to put this on my site just because the answers sounded good.

Before launching it, I tested it against questions I had actually been asked, as well as questions I expected someone visiting the site might ask. I ran each question multiple times to see where the answers broke down: where it made claims I couldn’t support, misunderstood what I was trying to say, or sounded more certain than I actually am.

The results below are from the final round, after I had already changed the assistant based on earlier testing.

What early testing caught

The assistant knew that I’m not deeply technical in areas like Linux, containers, or infrastructure.

That part was true.

But in some earlier testing, it could take a gap like that and jump too quickly to a much bigger conclusion about whether I would be a good fit for an infrastructure-heavy product role.

That wasn’t a fair representation of the evidence.

What changed

I tightened the instructions around how it talks about strengths, gaps, and fit.

If something is documented, it can say it. If it’s making an inference, it should make that clear. And a missing skill should not suddenly become a verdict on an entire role.

That sounds small, but it was exactly the kind of thing I wanted these tests to catch.

Final round of testing

After making those changes, I ran the full evaluation again for two reasons:

  • Make sure the changes held up consistently across repeated runs.
  • Compare the two models and see whether the more expensive model was actually better enough to justify using it.

18 questions · 6 areas · 3 runs each · 54 answers per model

All 18 questions cleared the release bar across three runs for both models.

Each answer needed to score at least 8/10, with no critical failure in grounding or boundaries.

ModelAnswers passedAverageCost
GPT-5.6 Sol54 / 5410.00$1.83
GPT-5.6 Luna54 / 549.98$0.07

Conclusion

Both models cleared the bar. There wasn’t enough meaningful difference to justify paying for the more expensive one, so I’m using Luna.

Sol sometimes wrote the answer differently. I couldn’t find evidence that it was giving me meaningfully better answers.

For this assistant, once the source material, instructions, and boundaries were in good shape, the more expensive model mostly gave me different prose—not a better result.

That was probably my biggest takeaway from building this: the model matters, but it is only one part of the product. A lot of the quality came from deciding what the assistant should know, what it should not claim, and what a good answer actually looks like.

Evaluation evidence

Inspect the evaluationQuestions, rubric, scores, and individual answers.

The test set contains 18 hiring-manager questions across six areas. Choose a question, then inspect one reviewed run at a time. The full 18 × 3 matrix remains in the stored evaluation data but is not the public story.

Current work

Differentiation

Contribution

Role fit

Technical judgment

Value creation

Current work · anchor

What has Jose been working on since leaving Zest AI?

Answer from

GPT-5.6 Luna · run 1

10/10 · pass

Grounding 2/2Directness 2/2Calibration 2/2Boundaries 2/2Usefulness 2/2

Since leaving Zest AI, Jose has been focused on turning his experience into independent products, writing, and a clearer next chapter of work.

  • Building AI products: He created Jose AI Assistant, a conversational website that lets visitors explore his work and judgment through grounded questions rather than a static portfolio.
  • Exploring personal decision systems: He has been developing Compass, a reflection and review system designed to help identify patterns, maintain commitments, and make better decisions without reducing life to checklists.
  • Building and testing small products: This includes Record Library, a music discovery and curation app; Capture, an exploration of surfacing useful knowledge at the right moment; and experiments involving agents, personal knowledge, and AI-assisted workflows.
  • Writing and refining his product philosophy: He has been developing ideas around AI adoption, trust, workflow design, human review, and why “the model is not the product.”
  • Clarifying his professional direction: He has been extracting the recurring patterns across finance, applied AI, and product management—especially how to turn complex systems into decisions people can understand, trust, and act on—and looking for roles where he can apply that judgment to consequential products.

The common thread is moving from private understanding to visible, testable work.

Scoring contract

Grounding

Claims can be traced to approved evidence.

Directness

The answer starts with the answer to the question.

Calibration

Fact, pattern, inference, and unknown stay distinct.

Boundaries

It avoids private, inflated, or unsupported claims.

Usefulness

A visitor leaves with a clearer picture or next question.

Five dimensions, scored 0–2. A response passes at 8/10 only when Grounding and Boundaries are nonzero. The tested model never grades itself: every answer was reviewed against a written answer key, and the model snapshot needed at least 80% passing answers.