Aletheion methodology evaluation · system_id: gpt-4o · audit date: 2026-03-23
| Dimension | Score | Tasks | Per-task scores | Result |
|---|---|---|---|---|
| perception | 0.993 | 2 (2p / 0f) | 1.00, 0.99 | PASS |
| generation | 0.717 | 2 (1p / 1f) | 0.45, 0.85 | CONDITIONAL |
| attention | 0.988 | 2 (2p / 0f) | 1.00, 0.98 | PASS |
| learning | 0.090 | 2 (0p / 2f) | 0.00, 0.15 | FAIL |
| memory | 0.990 | 1 (1p / 0f) | 0.99 | PASS |
| reasoning | 0.873 | 3 (3p / 0f) | 0.99, 0.90, 0.82 | PASS |
| metacognition | 0.747 | 2 (1p / 1f) | 0.94, 0.65 | CONDITIONAL |
| executive_functions | 0.740 | 1 (1p / 0f) | 0.74 | CONDITIONAL |
| problem_solving | 0.587 | 3 (1p / 2f) | 0.98, 0.50, 0.50 | CONDITIONAL |
| social_cognition | 0.900 | 1 (1p / 0f) | 0.90 | PASS |
| novelty | 0.713 | 3 (1p / 2f) | 0.65, 0.60, 0.82 | CONDITIONAL |
| orchestration | 0.433 | 2 (0p / 2f) | 0.50, 0.40 | FAIL |
corpus_taas_profiles[system_id=gpt-4o] · Audit definition: corpus_agi_audits[system_id=gpt-4o] · System profile: agi_evaluation_profiles[system=gpt-4o] · Persona runs: agi_complex_evaluations[system.system_id=gpt-4o] · Methodology config: agi_audits.config/api/v4/site-data/evaluation/gpt-4o/api/v4/site-data/methodology