# Difficult Problems Evaluation Sprint

A $7,500 founding-client professional evaluation service for one defined AI system, agent, workflow, or model implementation. It includes approximately 25 purpose-built difficult reasoning cases, baseline evaluation, structured failure analysis, a Reasoning Failure Report, remediation recommendations, and one re-test.

Canonical URL: https://newphysician.org/ai-evaluation/

## Related public resources
- [Sample Reasoning Failure Report](https://newphysician.org/ai-evaluation/sample-reasoning-failure-report/)
- [Start an Evaluation](https://newphysician.org/discuss/?interest=ai_evaluation)

Evidence note: In a 22-case local-model demonstration, two different models both failed 20 cases. Eight selected commercially meaningful failures were reviewed by Sonny Saggar and all eight were confirmed. This local demonstration does not make claims about all AI systems; GPT, Claude, and Gemini were not part of it.
