Skip to content

Data and ethics

This repository contains code only. It ships no collected YouTube data, no transcripts, no knowledge base and no trained models. The reasons come from primary sources collected in the publishing research.

The research is not legal advice. In short:

What Published? Why
Source code Yes Our own work; no credentials in it.
Video metadata and statistics No The YouTube API terms forbid redistributing API data, and data collected with an API key may be stored for at most 30 days.
Aggregated statistics, knowledge base No The policies also restrict creating and sharing derived metrics.
Transcripts No They reproduce the teachers' lectures, which the teachers own.
Trained models No Built from the data above; the terms do not clearly allow it, so we don't.

What the project collected

Channels 36 Algerian Bac channels, chosen by hand, one subject each
Videos discovered 19,919
Statistics fetched (2026-09-30) 19,042, for 435 API quota units
Bac 3AS lessons after filtering 9,801
Valid caption transcripts 4,583 collected, 4,381 for current Bac videos
Hand labels 300 videos labelled by hand to check the filter

How to reproduce

Bring your own YouTube Data API key and rebuild everything:

cp .env.example .env                        # set YOUTUBE_API_KEY
uv run python run_pipeline.py collect --channels config/channels.csv
uv run python run_pipeline.py filter_data
uv run python run_pipeline.py transcripts --input data/processed/videos_bac_only.csv --output data/processed/transcripts.csv
uv run python run_pipeline.py clean --transcripts data/processed/transcripts.csv
uv run python run_pipeline.py train

config/channels.csv is the list of channels to study, with columns channel_id,channel_name,subjects, and subjects must be one of the nine subjects. It holds the 36 public channels this project studied; being on the list is not an endorsement, and the file contains no statistics. Statistics refreshes are cheap: one quota unit per 50 videos.

Keep the 30-day limit in mind: delete or refresh API data within 30 days of collecting it.

Collecting responsibly

  • Quota. Collection uses the uploads playlist and batches of 50, about 100× fewer quota units than search.
  • Transcripts. Downloaded one video at a time with a pause between requests (--delay, default 2 s). After repeated refusals the collector waits; it never switches IP addresses.
  • People. No comments, no author names and no personal data are used for modelling. Channel and video IDs identify public educational channels.