Python Analytics

Nafs Results Analytics — Aseer Region

A complete Python pipeline that starts from the raw extract, applies data-quality rules and quarantines failures, then builds an analytical model of trends, gaps and school segments that feeds this executive dashboard.

  1. 1Raw extract70,041 rows
  2. 2DQ rulesDQ-001…007
  3. 3Modelpandas · SciPy
  4. 4DashboardJSON → Chart.js
Filters

Proficiency trend vs national

National benchmark weighted by the selected grade/subject mix

Key findings

Generated by the pipeline
  • Regional proficiency rose from 32.0% (1444 AH) to 40.4% (1447 AH), +8.4 pts at 2.9 pts per year.
  • Gap to the national average is -1.8 pts in 1447 AH vs -3.0 in 1444 AH.
  • Science leads (52.2%) and Mathematics trails (31.5%), a 20.7-pt spread; weakest sub-domain: Algebra (26.9%).
  • Top governorate Abha (45.3%), lowest Al-Harjah (29.5%); most improved Tanomah (+5.0).
  • Girls' schools outperform by 5.5 pts on average; the difference is statistically significant (Welch t-test, p < 0.001).
  • 189 schools flagged for priority intervention: ≥5 pts below the regional average and down ≥5 pts year-on-year (of 2,137 classified).
  • On the current trend, 1448 AH proficiency is projected at ~43.4% (80% range: 42.5–44.2%).

Proficiency by subject & grade

Performance-level distribution

Share of students per level

Governorate ranking

Dashed line: regional average

Gender & authority

Sub-domain diagnosis

Grade 6 science & Grade 9 maths

School segmentation: level vs change

Each dot is a school · 1447 AH vs 1446 · min. 15 tested

Priority schools for intervention

≥5 pts below average and down ≥5 pts · pseudonymised IDs
SchoolGovernorate1447Δ
SCH-1727 · Girls Ahad Rufaidah8.9%-22.4
SCH-0770 · Boys Bisha16.7%-23.9
SCH-1801 · Boys Bariq11.1%-16.2
SCH-1822 · Girls Al-Majardah20.7%-24.1
SCH-1592 · Boys Balqarn18.5%-21.5
SCH-1631 · Boys Balqarn10.3%-12.7
SCH-1337 · Girls Tathlith5.9%-7.9
SCH-2114 · Boys Al-Harjah8.0%-9.9
SCH-0016 · Boys Khamis Mushait9.1%-10.9
SCH-1020 · Girls Muhayil6.9%-7.4
SCH-1684 · Girls Ahad Rufaidah14.8%-15.2
SCH-1625 · Boys Balqarn14.8%-14.8

Total flagged: 189 · Full list in priority_schools.csv

Data quality before analysis

Raw-extract quality scorecard

The extract is checked against seven rules mapped to DAMA dimensions before any calculation. Failures are quarantined or conformed from reference data, with the action documented.

70,041
Raw rows
70,023
Rows certified
9
Quarantined
9
Duplicates removed
2
Values conformed
Completeness 99.99%
Consistency 100.00%
Validity 100.00%
Uniqueness 99.99%
Timeliness 100.00%
Scale 95–100% · dark tick = target
RuleDimensionCheckCheckedFailedPassAction
DQ-001CompletenessSchool ID is not null 70,041599.99%Quarantine & return to source
DQ-002ConsistencyGovernorate matches reference data 70,0412100.00%Auto-conform from reference table
DQ-003ValidityGrade & domain codes are in reference lists 70,0411100.00%Quarantine & fix coding
DQ-004ValidityProficiency within 0-100 and consistent with counts 70,041699.99%Recompute from source counts
DQ-005ValidityTested does not exceed expected 70,0413100.00%Quarantine for source verification
DQ-006UniquenessNo duplicate year/school/grade/domain 70,041999.99%Remove duplicates after review
DQ-007TimelinessExtract covers all reporting years 40100.00%Escalate delay to data owner
Method & statistics

Behind the numbers

Girls vs boys schools+5.5 pts Welch t = 8.58 · p < 0.001 · n = 1,112 / 1,025
Private & intl. vs public+4.6 pts Welch t = 4.07 · p < 0.001
School size vs proficiencyρ = -0.005 Spearman · p = 0.817 — no material relationship
Annual improvement+2.9 pts/yr OLS · R² = 0.996
Projection for 1448 AH43.4% 80% prediction interval: 42.5–44.2%

Pipeline excerpt

nafs_pipeline.py
# DQ-004 validity — proficiency within 0–100 and consistent with counts
recomputed = 100 * df["proficient"] / df["tested"]
bad_rate = (df["prof_rate"] < 0) | (df["prof_rate"] > 100) \
           | ((recomputed - df["prof_rate"]).abs() > 0.5)
rule("DQ-004", "Validity", bad_rate.sum(), action="Recompute from source counts")

# Priority = materially below average AND materially declining
wide["priority"] = ((wide["rate_1447"] <= region_avg - MATERIAL)
                    & (wide["change"] <= -MATERIAL)
                    & (wide["tested_1447"] >= PRIORITY_N))

# Significance: girls' vs boys' schools (Welch t-test)
t = stats.ttest_ind(girls, boys, equal_var=False)
  • DataSchool-level extract in the Nafs report layout: year, grade, domain & sub-domain, tested, proficient, performance levels and national average (1444–1447 AH).
  • PrivacySchool identifiers are pseudonymised; school-level indicators are suppressed below 15 tested students.
  • MeasuresProficiency = proficient ÷ tested (weighted by tested), because the score scale changed across years.

Static Python outputs for reports

Data note: figures on this page are generated from a modelled dataset that follows the Nafs report layout; the same pipeline runs unchanged on the official extract.