Skip to content

Проверять то, что курс о себе говорит — и сказать, что он есть по-английски - #13

Merged
DenisDrobyshev merged 4 commits into
mainfrom
add/social-cards-on-every-page
Sep 20, 2026
Merged

DenisDrobyshev merged 4 commits into
mainfrom
add/social-cards-on-every-page

Conversation

@DenisDrobyshev

@DenisDrobyshev DenisDrobyshev commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Курс много чего о себе утверждает, и ничто это не проверяет. Три файла и один шаблон — чтобы проверяло.

1. Каждой странице — карточку

Пятьдесят четыре страницы курса при отправке ссылки показывали голый адрес: ни заголовка, ни описания, ни картинки. Теги Open Graph были только на двух лендингах, вписанные руками, и картинки не было и там.

Это ровно наоборот тому, чем курс делятся. Делятся модулем — «вот модуль про дофамин и TD-обучение», — и именно у модуля показать было нечего.

overrides/main.html отдаёт теги каждой странице из её собственных title и meta.description. Отдельным блоком social_meta, чтобы лендинг мог их заменить, а не дописать: при дописывании og:title выходит дважды, и какой взять, решает сборщик карточки, а не мы. page пуста, пока тема рисует 404, поэтому каждое обращение закрыто проверкой — без этого сборка падает с 'None' has no attribute 'meta' и не говорит, в каком шаблоне.

Картинка — та же, что рисует генератор карточек организации: #0a0b0e, #f8f9fb, #d8b678 — чернила, бумага и золото самого сайта. twitter:card из summary становится summary_large_image: 1280×640, показанные миниатюрой, — потраченная впустую карточка.

2. Чтобы карточки не отвалились молча

Все способы сломать эти теги — молчаливые. Сборка проходит, --strict доволен, проверка ссылок слепа (теги — не ссылки), а глазами не увидеть: карточку видно только там, куда ссылку вставили, и не нам.

scripts/check_social.py читает собранный каталог и не пропускает семь случаев: тег пропал; тег пуст (так выглядит уронивший page 404 — заголовка нет на всём сайте); тег выведен дважды (лендинг дописал вместо замены); og:image относителен; og:image ведёт на файл, которого в сборке нет; og:url общий у двух страниц; twitter:card вернулся в summary.

Каждый из семи проверен тем, что копия сборки ломалась именно так и скрипт падал. Допуск «заголовок может повторяться дважды» намеренный: лендинги ru и en оба зовутся lemma.

3. Считать курс один раз, а не девять раз руками

Курс говорит «двадцать семь модулей в семи частях» в девяти местах — цифрами, русскими словами и английскими: бейдж, README дважды, обе программы, оба лендинга, CITATION.cff, .zenodo.json. Ни одно из девяти не знает, сколько модулей на самом деле. Добавить модуль — значит попасть в девять файлов и ни разу не ошибиться.

scripts/check_counts.py считает один раз — по docs/modules/, notebooks/ и таблицам программы — и держит девять мест при этом счёте. Ещё он складывает колонку «Время»: оттуда берутся оба итога программы, а сумму колонки никто не пересчитывает, правя одну строку.

Пропавшее утверждение — тоже расхождение. Перепишут фразу — шаблон перестанет находиться, проверка перестанет проверять и промолчит об этом. Это главный способ, которым такие тесты протухают, поэтому ненайденное место сообщается наравне с разошедшимся числом.

Что он нашёл на первом же прогоне

Части I–II названы как 8 недель. Их собственные таблицы дают 12: часть I — 1+2+2+2, часть II — 2+2+1. Соседний итог верен (таблицы дают 51, текст говорит «около 50»), так что 8 — описка, а не другой способ счёта, и она есть в обеих языковых версиях. Исправлено на 12 там и там. Если 8 значило что-то другое — это та строка, где надо возразить.

Второе нашлось в моём же скрипте: «в семь частях» вместо «в семи частях». У числительного есть падеж, и проверка, собирающая русскую фразу из числа, обязана его склонять, иначе падает на правильном тексте.

Проверено

  • mkdocs build --strict — проходит, 65 страниц. 404 собирается (это та страница, на которой сборка падала).
  • У всех 65 свой og:title, ни одного дубля, картинка абсолютна и лежит в сборке.
  • Четырнадцать способов разойтись проверены на копии: бейдж, каждый лендинг, каждый файл цитирования, переписанная фраза, модуль, выпавший из навигации, удалённый ноутбук, непереведённая страница, задвоенный номер, изменившееся время, разошедшиеся языки и неразбираемый .zenodo.json. Тринадцать ловились как написано; четырнадцатый падал трейсбэком вместо строки — теперь строкой.
  • Оба скрипта — только стандартная библиотека. Ставить в CI нечего.

4. Сказать на первой странице, что курс есть и по-английски

Сайт двуязычен давно — двадцать семь модулей по-русски и двадцать семь по-английски, за каждым один и тот же ноутбук, — а репозиторий об этом молчал. На GitHub была только русская страница, и англоязычный пришедший видел русский текст про курс, который он на самом деле может читать.

README.en.md — полный README, а не заглушка со ссылкой на сайт: попавший на GitHub решает там, по бейджам, структуре и таблице готовности, и заглушка просит его кликнуть раньше, чем он поймёт, стоит ли. В обоих файлах появилась строка переключения языка, та же, что в остальных репозиториях.

Два утверждения были не пропущены, а устарели. CITATION.cff и .zenodo.json говорили «Written in Russian» — это перестало быть правдой, когда английская сборка была доделана. Теперь оба говорят про обе версии и несут ключевое слово. language: rus в Zenodo оставлен: поле однозначное, русский — исходный язык. Если индексировать надо английскую сторону — это то место, где надо сказать.

languages в CITATION.cff не добавлен, хотя выглядит как ровно то поле. В CFF 1.2.0 languages существует только внутри definitions.reference, а не на верхнем уровне, и незнакомый верхнеуровневый ключ — это способ тихо погасить кнопку «Cite this repository».

Английский README иначе стал бы десятой поверхностью, утверждающей «двадцать семь модулей в семи частях» без всякой связи со счётом — ровно та проблема, про которую предыдущий пункт, — поэтому его бейдж и два словесных утверждения проверяются наравне с остальными.

check_counts.py заодно перестал говорить «девять мест» и считает их сам: сегодня шестнадцать утверждений в восьми файлах. Скрипту, вся работа которого в том, чтобы никто не держал число в голове, не годится держать число в голове.

Что проверено локально

  54 страниц, 27 ноутбуков — все связаны и открываются в Colab.
  27 модулей в 7 частях, 27 ноутбуков, 51 неделя — и все 16 мест, где курс это говорит (8 файлов), говорят то же.
  65 страниц, у каждой свой заголовок и адрес, картинка одна и она на месте.

плюс mkdocs build --strict, .zenodo.json разбирается, CITATION.cff разбирается и сохраняет свои девять ключей.

Only the two landings carried Open Graph tags, hand-written, and neither had an
image. The other fifty-four pages had none at all: shared anywhere, a module
produced a bare link with no title and no card.

That is the wrong way round. The shareable unit of a course is a module -- "вот
модуль про дофамин и TD-обучение" -- and the module was the one thing with
nothing to show.

main.html emits the tags for every page from its own title and description, in
a `social_meta` block so the landings can replace them rather than append:
appending would have emitted og:title twice and left the scraper to choose.
`page` is None while the theme renders the 404, so every access is guarded --
without that the build dies with "'None' has no attribute 'meta'" and does not
say which template.

The card is the one the organisation's generator already draws for lemma. Its
palette is not a coincidence: #0a0b0e, #f8f9fb and #d8b678 are the site's own
ink, paper and gold, so the card and the page a reader lands on are the same
design.

twitter:card goes from summary to summary_large_image, because a 1280×640 card
shown as a thumbnail is a wasted card.
The tags from the previous commit come out of a template, not out of the pages,
and every way they break is a way the build still passes.

`page` is empty while the theme renders the 404. Jinja does not raise there --
it substitutes an empty string -- so one unguarded access takes og:title off
all sixty-five pages and `mkdocs build --strict` reports success. Rename the
card and og:image keeps pointing at the old name. Move `social_meta` so a
landing appends instead of replacing and og:title comes out twice, leaving the
scraper to choose. None of this is a link, so the link checker is blind to it,
and none of it is visible on the site itself: a card is only ever seen
somewhere else, by someone who is not us.

check_social.py reads the built directory and refuses all seven: a missing,
empty or duplicated tag, an og:image that is relative or points at a file the
build does not contain, an og:url shared by two pages, twitter:card downgraded
from summary_large_image, and one title repeated across the site -- which is
what "the template stopped seeing page" looks like from the outside.

Each of the seven was checked by breaking a copy of the build and confirming
the script fails. The two-page allowance on repeated titles is deliberate: the
ru and en landings are both called lemma, and that is correct.

Standard library only, and it runs against the directory mkdocs has just
built, so the site job gains a step and nothing to install.
lemma says "двадцать семь модулей в семи частях" in nine places -- in digits,
in Russian words and in English words: the badge, README prose twice, both
programmes, both landings, CITATION.cff and .zenodo.json. None of the nine
knows how many modules there are. Adding one means hitting nine files and never
slipping, and the first slip survives until somebody counts by hand.

check_counts.py counts once -- docs/modules/, notebooks/, and the programme's
own tables -- and holds the nine to it. It also sums the "Время" column, which
is where the two totals come from: "около 50 недель" and the figure given for
parts I-II. Nobody re-adds a column while editing one row.

A claim that has gone missing counts as a discrepancy too. Rewrite the sentence
and the pattern stops matching, the check stops checking, and it says nothing
about that -- which is the failure mode of this whole kind of test. So an
unfound claim is reported exactly like a wrong number.

It found two things on its first run.

Parts I-II were described as 8 weeks. Their own tables give 12: part I is
1+2+2+2, part II is 2+2+1. The neighbouring total is right -- the tables sum to
51 and the text says about 50 -- so 8 is a slip, not a different way of
counting, and the same slip is in both languages. Changed to 12 in both. If 8
was meant as something else, this is the line to push back on.

The second was mine: "в семь частях" instead of "в семи частях". A numeral has
a case, and a check that builds Russian out of a number has to decline it, or
it fails on correct text. RU_PREP_IRREGULAR names the forms the rule misses --
including восьми, which is not восеми.

Fourteen ways of drifting were tried against a copy: badge, each landing, each
citation file, a rewritten sentence, a module dropped from the navigation, a
deleted notebook, a missing translation, a duplicated number, a changed
duration, the two languages disagreeing, and .zenodo.json left unparseable.
Thirteen were caught as written; the fourteenth raised instead of reporting,
and now reports.

Standard library, no build needed.
@DenisDrobyshev DenisDrobyshev changed the title Дать каждой странице карточку — и не дать ей отвалиться молча Проверять то, что курс о себе говорит: карточки, ссылки, числа Sep 20, 2026
The site has been bilingual for a while -- twenty-seven modules in Russian and
twenty-seven in English, the same notebook behind each -- and the repository
never said so. GitHub's front page was Russian only, so an English-speaking
arrival saw a Russian page about a course they can in fact read.

README.en.md is the whole README, not a stub pointing at the site: someone who
lands on GitHub decides there, from the badges, the layout and the table of
what is done, and a stub asks them to click before they know whether it is
worth clicking. Both files now carry the language line the other repositories
use.

Two claims were out of date rather than missing. CITATION.cff and .zenodo.json
both said "Written in Russian", which stopped being true when the English build
was finished; both now say Russian and English, both complete, and carry the
keyword. .zenodo.json keeps `language: rus` -- the field takes one value and
Russian is the source language. Flip it if the English side should be the one
indexed.

CITATION.cff gets no `languages` key. It reads like the place for this, but in
CFF 1.2.0 `languages` exists only inside `definitions.reference`, not at the
top level, and an unknown top-level key is how "Cite this repository" quietly
stops rendering.

README.en.md would otherwise be a tenth surface asserting "twenty-seven modules
in seven parts" with nothing holding it to the count -- the exact problem the
previous commit is about -- so its badge and its two spelled-out claims are
checked like the rest.

check_counts.py also stops saying "nine places". It counts them: sixteen
assertions across eight files today. A script whose whole job is that nobody
should hold a number in their head had no business holding one.
@DenisDrobyshev DenisDrobyshev changed the title Проверять то, что курс о себе говорит: карточки, ссылки, числа Проверять то, что курс о себе говорит — и сказать, что он есть по-английски Sep 20, 2026
@DenisDrobyshev
DenisDrobyshev merged commit c8b7887 into main Sep 20, 2026
5 checks passed
@DenisDrobyshev
DenisDrobyshev deleted the add/social-cards-on-every-page branch September 20, 2026 15:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant