NeoGen
An ongoing side research project on drug interaction: given a molecular structure, predict how a compound is absorbed and cleared, and how two drugs taken together change each other's behaviour. Machine-learned property models feed a simulation of the body over time, so a prediction comes with a mechanism rather than a label.
Benchmark v1.2.0: 79.4% of predictions land within 2-fold of published values, and roughly 80% within 2× across 30 clinically studied drug pairs. The harness prints its own failures — theophylline, atorvastatin, metoprolol. Evaluation design reviewed with Dr. Joga Gobburu, former Director of the FDA Division of Pharmacometrics.
The benchmark harness prints its own misses. Worst offenders in v1.2.0: theophylline, atorvastatin, metoprolol. They stay in the report because a benchmark that hides its failures is marketing.
The Scorecard
| All measures within 2-fold of published values | 81/102 · 79.4% |
| Drug-behaviour benchmark within 2× | 56/72 · 77.8% |
| Average fold error | 1.65× |
| Peak concentration within 2× | 79.2% |
| Total exposure within 2× | 83.3% |
| Clearance half-life within 2× | 70.8% |
| Acceptance gate (≥60% within 2×) | PASS |