Skip to content

Engagement model card

What it predicts

The engagement score of a Bac 3AS lesson on YouTube:

engagement score = log(1 + (3 × comments + likes) / √views)

Comments count three times as much as likes because on lesson videos they are mostly students asking questions. Dividing by the square root of views keeps large channels from dominating, and the log tames the long tail.

The predictor also returns a category: Low, Medium or High, meaning the bottom, middle or top third of training videos. The cut-offs are learned when the model is trained (2.106 and 2.764 for the current model).

Intended use

Scoring a planned video, before it is published, so a creator can compare versions of a title, description or length. It is an estimate for one niche (Algerian Bac 3AS lessons), not a guarantee of views.

Inputs

Field Required Notes
title yes
duration_sec yes
channel_id or subject one of them An unseen channel gets the training medians for channel features, and the result says known_channel: false.
description, tags, publish_date, transcript_text no A missing date counts as "published now".

No statistics are needed, and none are read: a test checks that the features are identical with view, like and comment counts present, set to NaN, or absent.

Model

  • Algorithm: random forest (357 trees, max depth 21, max_features=0.5). These are the best settings from the notebook's Optuna search.
  • 171 features:
    • 41 numeric features, kept by VIF selection from 55 candidates. The candidates are title and description lengths, exam and pedagogical keywords, duration, channel statistics, subject, and 32 transcript features such as readability, speaking pace and questions asked
    • 30 TF-IDF terms
    • 100 PCA components of AraBERT v2 embeddings (aubmindlab/bert-base-arabertv2) of title + description
  • Fitting: everything data-dependent is fit on the training split only: channel statistics, VIF selection, scaler, TF-IDF and PCA.

Training data

Collected 2026-09-30, YouTube Data API v3
Channels 36 curated, 35 still available
Videos 9,801 Bac 3AS lessons (filtered from 19,042)
With a valid transcript 4,381
Published 2014-03-15 to 2026-06-06
Subjects Maths 4,097 · History & Geography 1,910 · Natural Sciences 1,710 · Physics 894 · Islamic Sciences 369 · Arabic 352 · English 236 · French 124 · Philosophy 109

Results

Held-out test set: the 1,963 videos of the notebook's test split that are still online (random 20%, seed 42). Training used the other 7,838.

Metric Value
R² 0.703
MAE 0.314
RMSE 0.410

How it compares to the notebook

The exploration notebook reported R² 0.693 for the same algorithm. That number is optimistic: the notebook fitted channel statistics, feature selection and the scaler on all videos, test videos included. The production pipeline fixes that leak. The data also differs: statistics are eight months more mature and 52 test videos are gone. The two numbers therefore measure different things, and the drop from leak-free fitting is not separately measured.

The notebook's comparison of algorithms (January data, same leaky protocol, useful for ranking only):

Model MAE RMSE R²
Random Forest (Optuna) 0.310 0.413 0.692
Ensemble (all 4 models) 0.312 0.414 0.690
CatBoost (Optuna) 0.313 0.418 0.684
Random Forest 0.317 0.422 0.678
Neural network (MLP) 0.316 0.426 0.672
XGBoost (regularized) 0.330 0.434 0.660
LightGBM 0.337 0.441 0.649
XGBoost 0.337 0.443 0.645
Ridge regression 0.416 0.530 0.492
Lasso regression 0.599 0.744 ≈ 0

Limitations

  • Same channels in train and test. The split is random over videos, so every test video's channel is also in training. How well the model does on a channel it has never seen is not measured, and that is the coach's main use case.
  • The score is a proxy. It measures visible interaction, not learning.
  • Snapshot in time. Older videos have had longer to collect comments and likes; the model has no "age" feature, by design, because a planned video has no age.
  • Uneven subjects. Maths is 42% of the data and Philosophy 1%; predictions for small subjects rest on fewer examples.
  • Three language subjects have no curriculum keyword list. Their keyword features use all subjects' keywords.

Reproducing

The model is trained from data that is not published (see Data and ethics). With your own collection:

uv run python run_pipeline.py train --split-file splits_mapping.csv   # or omit for a random 20%