DeepSeek-R1 Competent

Aletheion methodology evaluation · system_id: deepseek-r1 · audit date: 2026-03-23

78.2
Vendor: DeepSeek
Tier: Competent · rank 51
Tasks: 24 total · 19 passed · 5 failed
Methodology version: 1.0.0
Composite score: 0.782 PASS

Cognitive-dimension radar

PerceptionGenerationAttentionLearningMemoryReasoningMetacognitionExecutive functionsProblem solvingSocial cognitionNoveltyOrchestration

Dimension scores

Perception100.0Generation81.7Attention99.2Learning57.0Memory98.0Reasoning90.5Metacognition68.0Executive functions87.0Problem solving60.4Social cognition90.0Novelty72.5Orchestration60.7

Per-dimension breakdown (task-level)

DimensionScoreTasksPer-task scoresResult
perception1.000 2 (2p / 0f)1.00, 1.00PASS
generation0.817 2 (2p / 0f)0.85, 0.80PASS
attention0.992 2 (2p / 0f)0.98, 1.00PASS
learning0.570 2 (1p / 1f)0.00, 0.95CONDITIONAL
memory0.980 1 (1p / 0f)0.98PASS
reasoning0.904 3 (3p / 0f)1.00, 0.95, 0.85PASS
metacognition0.680 2 (1p / 1f)0.92, 0.56CONDITIONAL
executive_functions0.870 1 (1p / 0f)0.87PASS
problem_solving0.604 3 (2p / 1f)0.98, 0.20, 0.68CONDITIONAL
social_cognition0.900 1 (1p / 0f)0.90PASS
novelty0.725 3 (2p / 1f)0.65, 0.70, 0.78CONDITIONAL
orchestration0.607 2 (1p / 1f)0.70, 0.56CONDITIONAL

Strongest dimensions

Weakest dimensions

Evidence provenance
Profile: corpus_taas_profiles[system_id=deepseek-r1] · Audit definition: corpus_agi_audits[system_id=deepseek-r1] · System profile: agi_evaluation_profiles[system=deepseek-r1] · Persona runs: agi_complex_evaluations[system.system_id=deepseek-r1] · Methodology config: agi_audits.config
Per-dimension evidence: /evidence/alitheion-eval-deepseek-r1

Live API endpoint: /api/v4/site-data/evaluation/deepseek-r1
Methodology document: /api/v4/site-data/methodology