Section topics

The section topics data pipeline is a sequence of pyspark.sql.DataFrame extraction and transformation functions.

Inputs come from Wikimedia Foundation’s Analytics Data Lake:

High-level steps:

  • gather wikitext of sections at a given hierarchy level via the MediaWiki parser from hell. Default: section_topics.pipeline.SECTION_LEVEL, lead section included

  • optionally filter out sections that don’t convey relevant content, typically lists and tables

  • extract Wikidata QIDs from wikilinks: the so-called section topics

  • optionally filter out noisy topics, typically dates and numbers

  • compute the relevance score

Output row example:

snapshot

wiki_db

page_namespace

revision_id

page_qid

page_id

page_title

section_index

section_title

topic_qid

topic_title

topic_score

2023-01-16

enwiki

0

1127523670

Q36724

841

Attila

5

Solitary kingship

Q3623581

Arnegisclus

1.13

More documentation lives in MediaWiki.

Functions are ordered by their execution in the pipeline.

section_topics.pipeline.SECTION_LEVEL = 2

Section hierarchy level to be extracted. Level 1 is just for page titles, actual sections start from level 2.

section_topics.pipeline.SECTION_ZERO_HEADING = '### zero ###'

Reserved heading for the lead section, AKA section zero.

section_topics.pipeline.STRIP_CHARS = '!"#$%&\' *+,-./:;<=>?@[\\]^_`{|}~'

ASCII punctuation characters to be stripped from section headings. Include the ASCII white space, don’t strip round brackets.

section_topics.pipeline.SUBSTITUTE_PATTERN = '[\\s_]'

All kinds of white space to be substituted for the ASCII one; underscores turn into spaces as well.

section_topics.pipeline.RECOGNIZED_HTML_TAGS = ['tt', 'wbr', 'br', 'sup', 'var', 's', 'ruby', 'sub', 'ul', 'abbr', 'hr', 'p', 'li', 'ins', 'cite', 'pre', 'small', 'bdi', 'h6', 'em', 'h1', 'rp', 'table', 'strike', 'data', 'blockquote', 'i', 'div', 'mark', 'center', 'dd', 'tr', 'h4', 'u', 'rt', 'span', 'link', 'td', 'dt', 'th', 'strong', 'caption', 'big', 'rtc', 'dfn', 'time', 'code', 'dl', 'del', 'samp', 'b', 'font', 'rb', 'h5', 'bdo', 'ol', 'kbd', 'q', 'h2', 'meta', 'h3']

Allowed HTML tags to be stripped from section headings. Based on MediaWiki’s Sanitizer

section_topics.pipeline.LIST_OR_TABLE_RE = re.compile('^\\*|^#|\\{\\|', re.MULTILINE)

Match standard wikitext lists and tables:

  • line starts with * -> unordered list

  • line starts with # -> ordered list + avoid link anchors

  • {| anywhere -> table

section_topics.pipeline.MINIMUM_SECTION_CHARS = 500

Section content length threshold

section_topics.pipeline.load_articles(spark, wikis)[source]

Load the current revision of all articles (redirects included) through the section_topics.queries.ARTICLES Data Lake query.

Parameters:
  • spark (SparkSession) – an active Spark session

  • wikis (str) – a string of comma-separated wikis to process

Return type:

DataFrame

Returns:

the dataframe of all articles

section_topics.pipeline.load_qids(spark, weekly_snapshot, wikis)[source]
Load Wikidata QIDs with their page links through the

section_topics.queries.QIDS Data Lake query.

Parameters:
  • spark (SparkSession) – an active Spark session

  • weekly_snapshot (str) – a YYYY-MM-DD date

  • wikis (str) – a string of comma-separated wikis to process

Return type:

DataFrame

Returns:

the dataframe of QIDs and page links

section_topics.pipeline.separate_redirects(articles_all)[source]

Separate redirects from actual articles.

Parameters:

articles_all (DataFrame) – a dataframe of all articles as returned by load_articles()

Return type:

tuple

Returns:

the dataframes of articles and redirects

section_topics.pipeline.apply_filter(df, filter_df, broadcast=False)[source]

Exclude rows of an input dataframe given a filter dataframe.

Anti-join all filter columns against input ones.

Parameters:
  • df (DataFrame) – a dataframe to be filtered

  • filter_df (DataFrame) – a dataframe acting as a filter. Columns must be a subset of df

  • broadcast (bool) – whether to broadcast filter_df, which tells Spark to perform a broadcast hash join, i.e., pyspark.sql.functions.broadcast(). Much faster if filter_df is small

Return type:

DataFrame

Returns:

the filtered df dataframe

section_topics.pipeline.look_up_qids(pages, qids)[source]

Look up page QIDs through page titles.

Parameters:
  • pages (DataFrame) – a dataframe of pages as output by apply_filter(). Pass the output of load_pages() if you want the full raw dataset.

  • qids (DataFrame) – a dataframe of page IDs and Wikidata QIDs as output by load_qids()

Return type:

DataFrame

Returns:

the pages dataframe with page QIDs added

section_topics.pipeline.wikitext_headings_to_anchors(headings)[source]

Transform wikitext headings into URL anchors.

For instance, === Album in studio === becomes Album_in_studio, and serves as a section link in https://it.wikipedia.org/wiki/Gaznevada#Album_in_studio.

Anchors that occur more than once get a numeric suffix in the form anchor_N.

Parameters:

headings (List[str]) – a list of wikitext headings

Return type:

List[str]

Returns:

the corresponding URL anchors

section_topics.pipeline.normalize_heading_column(column, substitute_pattern='[\\\\s_]', strip_chars='!"#$%&\\' *+, -./:;<=>?@[\\\\]^_`{|}~')[source]

Normalize a dataframe column of section headings for better matching.

Normalization steps:

  • remove _N suffixes in case of duplicate section anchors as added by wikitext_headings_to_anchors()

  • replace characters matched by substitute_re with one ASCII white space

  • strip leading and trailing strip_chars

  • lowercase

Note

This normalization is not perfect: it’s a trade-off between several ones, some of which may prevent from converging to a lowest common denominator. However, only extreme edge cases might be affected.

Parameters:
  • column (str) – a dataframe column name of section headings

  • substitute_pattern (str) – (optional) a regular expression pattern whose matches will be replaced by one ASCII white space

  • strip_chars (str) – (optional) a string of characters to be stripped

Return type:

Column

Returns:

the column of normalized section headings

section_topics.pipeline.parse_wikitext(keep_lists_and_tables, minimum_section_size)[source]

Currying function that passes args to the underlying parse() PySpark user-defined function (UDF).

See also this gist.

Parameters:
  • keep_lists_and_tables (bool) – whether to keep sections with at least one standard wikitext list or table

  • minimum_section_size (int) – minimum content character length for the section to be considered

Return type:

udf

Returns:

the actual UDF

section_topics.pipeline.extract_sections(articles, keep_lists_and_tables, minimum_section_size)[source]

Extract section data from articles

Parameters:
  • articles (DataFrame) – a dataframe of article pages as output by look_up_qids()

  • keep_lists_and_tables (bool) – whether to keep sections with at least one standard wikitext list or table

  • minimum_section_size (int) – minimum content character length for the section to be considered

Return type:

DataFrame

Returns:

DF of sections with links (1 row per link), DF of sections with images (1 row per section)

Lowercase the first character of wikilink target titles.

The link target is case-sensitive except for the first character.

Parameters:

link_column (str) – a dataframe column name of wikilinks

Return type:

Column

Returns:

the column of normalized wikilinks. None values are kept.

section_topics.pipeline.handle_media(sections, media_prefixes)[source]

Separate media links from other ones.

Detect media links via lowercased lookup of namespace prefixes.

Parameters:
  • sections (DataFrame) – a dataframe of sections and wikilinks as output by extract_sections()

  • media_prefixes (list) – a list of namespace prefixes for media pages

Return type:

Tuple[DataFrame, DataFrame]

Returns:

the dataframe of media links and the remainder of the sections dataframe

Filter empty strings and add a column of normalized wikilinks.

Parameters:

sections (DataFrame) – a dataframe of sections and wikilinks as output by extract_sections()

Return type:

DataFrame

Returns:

the cleaned sections dataframe

section_topics.pipeline.resolve_redirects(sections, redirects)[source]

Follow section wikilinks redirects.

If a wikilink points to a redirect page, replace its title with the redirected one. Normalize redirects via normalize_wikilinks().

Parameters:
  • sections (DataFrame) – a dataframe of sections and cleaned wikilinks as output by clean_up_links()

  • redirects (DataFrame) – a dataframe of page titles and redirected page titles as output by load_redirects()

Return type:

DataFrame

Returns:

the dataframe of redirected wikilinks. Both original and normalized ones are kept.

section_topics.pipeline.gather_section_topics(sections, articles)[source]

Align section wikilinks to their Wikidata QIDs: the so-called section topics.

Parameters:
  • sections (DataFrame) – a dataframe of sections and normalized wikilinks as output by resolve_redirects(). Pass the output of clean_up_links() to skip wikilinks pointing to redirect pages.

  • articles (DataFrame) – a dataframe of articles as output by look_up_qids()

Return type:

DataFrame

Returns:

the dataframe of sections and topics (as Wikidata QIDs). Original topic titles are kept.

section_topics.pipeline.compute_relevance(topics, level='section')[source]

Compute either the section-level or the article-level relevance score for every section topic.

The section-level score is a standard term frequency-inverted document frequency (TF-IDF). The article-level score is a custom TF-IDF, where TF is across wikis and IDF is within one wiki.

Workflow:

  1. filter null topic QIDs

  2. compute TF numerator: occurrences of one topic QID in a section or page QID

  3. compute TF denominator: occurrences of all topic QIDs in a section or page QID

  4. compute TF: numerator / denominator

  5. join with input on section or page QID and topic QID

  6. compute IDF numerator: count of sections or page QIDs in a wiki

  7. compute IDF denominator: count of sections or page QIDs where a topic QID occurs, in a wiki

  8. compute IDF: log( numerator / denominator )

  9. join with input on wiki and topic QID

  10. compute TF-IDF: TF * IDF

Parameters:
  • topics (DataFrame) – a dataframe of section topics as output by gather_section_topics()

  • level (str) – (optional) at which level relevance is computed, section or article

Return type:

DataFrame

Returns:

the input dataframe with the tf_idf column added

section_topics.pipeline.compose_output(scored_topics, all_topics, snapshot, page_namespace=0)[source]

Fuse scored topics with null ones and build the output dataset.

Parameters:
  • scored_topics (DataFrame) – a dataframe of scored topics as output by compute_relevance()

  • all_topics (DataFrame) – a dataframe of all topics as output by gather_section_topics()

  • snapshot (str) – a weekly snapshot to serve as the constant value for the snapshot column of the output dataframe

  • page_namespace (int) – (optional) a page namespace to serve as the constant value for the page_namespace column of the output dataframe

Return type:

DataFrame

Returns:

the final output dataframe

section_topics.pipeline.parse(wikitext, section_level=SECTION_LEVEL, section_zero_heading=SECTION_ZERO_HEADING)

Note

This is the core function responsible for the data heavy lifting. It’s implemented as a PySpark user-defined function (pyspark.sql.functions.udf()). It must be called by the parse_wikitext() currying function, which seems the only way to pass an optional denylist of section headings.

Parse a wikitext string into a list of dicts. Each dict represents a section within the given wikitext and has the following keys:

  • index - section number, 0 for the leading section

  • heading - section heading, section_zero_heading’s value for the leading section

  • links - list of section wikilinks

  • images - list of section image file names

A section heading is normalized through wikitext_headings_to_anchors(). Clean wikilinks with mwparserfromhell’s strip_code(). This can lead to empty strings that should be filtered.

Parameters:
  • wikitext (str) – a wikitext

  • section_level (int) – (optional) a section hierarchy level to be extracted

  • section_zero_heading (str) – (optional) a heading reserved to the lead section

Returns:

the list of (index, heading, links, images) section dictionaries

Return type:

List[str]