Aletheion methodology evaluation · system_id: deepseek-r1 · audit date: 2026-03-23
| Dimension | Score | Tasks | Per-task scores | Result |
|---|---|---|---|---|
| perception | 1.000 | 2 (2p / 0f) | 1.00, 1.00 | PASS |
| generation | 0.817 | 2 (2p / 0f) | 0.85, 0.80 | PASS |
| attention | 0.992 | 2 (2p / 0f) | 0.98, 1.00 | PASS |
| learning | 0.570 | 2 (1p / 1f) | 0.00, 0.95 | CONDITIONAL |
| memory | 0.980 | 1 (1p / 0f) | 0.98 | PASS |
| reasoning | 0.904 | 3 (3p / 0f) | 1.00, 0.95, 0.85 | PASS |
| metacognition | 0.680 | 2 (1p / 1f) | 0.92, 0.56 | CONDITIONAL |
| executive_functions | 0.870 | 1 (1p / 0f) | 0.87 | PASS |
| problem_solving | 0.604 | 3 (2p / 1f) | 0.98, 0.20, 0.68 | CONDITIONAL |
| social_cognition | 0.900 | 1 (1p / 0f) | 0.90 | PASS |
| novelty | 0.725 | 3 (2p / 1f) | 0.65, 0.70, 0.78 | CONDITIONAL |
| orchestration | 0.607 | 2 (1p / 1f) | 0.70, 0.56 | CONDITIONAL |
corpus_taas_profiles[system_id=deepseek-r1] · Audit definition: corpus_agi_audits[system_id=deepseek-r1] · System profile: agi_evaluation_profiles[system=deepseek-r1] · Persona runs: agi_complex_evaluations[system.system_id=deepseek-r1] · Methodology config: agi_audits.config/api/v4/site-data/evaluation/deepseek-r1/api/v4/site-data/methodology