Section topics¶
The section topics data pipeline is a sequence of pyspark.sql.DataFrame extraction and transformation functions.
Inputs come from Wikimedia Foundation’s Analytics Data Lake:
Wikipedias wikitext (all Wikipedias by default as per wikipedias.txt )
High-level steps:
gather wikitext of sections at a given hierarchy level via the MediaWiki parser from hell. Default:
section_topics.pipeline.SECTION_LEVEL, lead section includedoptionally filter out sections that don’t convey relevant content, typically lists and tables
extract Wikidata QIDs from wikilinks: the so-called section topics
optionally filter out noisy topics, typically dates and numbers
compute the relevance score
Output row example:
snapshot |
wiki_db |
page_namespace |
revision_id |
page_qid |
page_id |
page_title |
section_index |
section_title |
topic_qid |
topic_title |
topic_score |
|---|---|---|---|---|---|---|---|---|---|---|---|
2023-01-16 |
enwiki |
0 |
1127523670 |
Q36724 |
841 |
Attila |
5 |
Solitary kingship |
Q3623581 |
Arnegisclus |
1.13 |
More documentation lives in MediaWiki.
Functions are ordered by their execution in the pipeline.
- section_topics.pipeline.SECTION_LEVEL = 2¶
Section hierarchy level to be extracted. Level 1 is just for page titles, actual sections start from level 2.
- section_topics.pipeline.SECTION_ZERO_HEADING = '### zero ###'¶
Reserved heading for the lead section, AKA section zero.
- section_topics.pipeline.STRIP_CHARS = '!"#$%&\' *+,-./:;<=>?@[\\]^_`{|}~'¶
ASCII punctuation characters to be stripped from section headings. Include the ASCII white space, don’t strip round brackets.
- section_topics.pipeline.SUBSTITUTE_PATTERN = '[\\s_]'¶
All kinds of white space to be substituted for the ASCII one; underscores turn into spaces as well.
- section_topics.pipeline.RECOGNIZED_HTML_TAGS = ['tt', 'wbr', 'br', 'sup', 'var', 's', 'ruby', 'sub', 'ul', 'abbr', 'hr', 'p', 'li', 'ins', 'cite', 'pre', 'small', 'bdi', 'h6', 'em', 'h1', 'rp', 'table', 'strike', 'data', 'blockquote', 'i', 'div', 'mark', 'center', 'dd', 'tr', 'h4', 'u', 'rt', 'span', 'link', 'td', 'dt', 'th', 'strong', 'caption', 'big', 'rtc', 'dfn', 'time', 'code', 'dl', 'del', 'samp', 'b', 'font', 'rb', 'h5', 'bdo', 'ol', 'kbd', 'q', 'h2', 'meta', 'h3']¶
Allowed HTML tags to be stripped from section headings. Based on MediaWiki’s Sanitizer
- section_topics.pipeline.LIST_OR_TABLE_RE = re.compile('^\\*|^#|\\{\\|', re.MULTILINE)¶
Match standard wikitext lists and tables:
line starts with
*-> unordered listline starts with
#-> ordered list + avoid link anchors{|anywhere -> table
- section_topics.pipeline.MINIMUM_SECTION_CHARS = 500¶
Section content length threshold
- section_topics.pipeline.load_articles(spark, wikis)[source]¶
Load the current revision of all articles (redirects included) through the
section_topics.queries.ARTICLESData Lake query.- Parameters:
spark (
SparkSession) – an active Spark sessionwikis (
str) – a string of comma-separated wikis to process
- Return type:
DataFrame- Returns:
the dataframe of all articles
- section_topics.pipeline.load_qids(spark, weekly_snapshot, wikis)[source]¶
- Load Wikidata QIDs with their page links through the
section_topics.queries.QIDSData Lake query.
- section_topics.pipeline.separate_redirects(articles_all)[source]¶
Separate redirects from actual articles.
- Parameters:
articles_all (
DataFrame) – a dataframe of all articles as returned byload_articles()- Return type:
- Returns:
the dataframes of articles and redirects
- section_topics.pipeline.apply_filter(df, filter_df, broadcast=False)[source]¶
Exclude rows of an input dataframe given a filter dataframe.
Anti-join all filter columns against input ones.
- Parameters:
df (
DataFrame) – a dataframe to be filteredfilter_df (
DataFrame) – a dataframe acting as a filter. Columns must be a subset ofdfbroadcast (
bool) – whether to broadcastfilter_df, which tells Spark to perform a broadcast hash join, i.e.,pyspark.sql.functions.broadcast(). Much faster iffilter_dfis small
- Return type:
DataFrame- Returns:
the filtered
dfdataframe
- section_topics.pipeline.look_up_qids(pages, qids)[source]¶
Look up page QIDs through page titles.
- Parameters:
pages (
DataFrame) – a dataframe of pages as output byapply_filter(). Pass the output ofload_pages()if you want the full raw dataset.qids (
DataFrame) – a dataframe of page IDs and Wikidata QIDs as output byload_qids()
- Return type:
DataFrame- Returns:
the
pagesdataframe with page QIDs added
- section_topics.pipeline.wikitext_headings_to_anchors(headings)[source]¶
Transform wikitext headings into URL anchors.
For instance,
=== Album in studio ===becomesAlbum_in_studio, and serves as a section link in https://it.wikipedia.org/wiki/Gaznevada#Album_in_studio.Anchors that occur more than once get a numeric suffix in the form
anchor_N.
- section_topics.pipeline.normalize_heading_column(column, substitute_pattern='[\\\\s_]', strip_chars='!"#$%&\\' *+, -./:;<=>?@[\\\\]^_`{|}~')[source]¶
Normalize a dataframe column of section headings for better matching.
Normalization steps:
remove
_Nsuffixes in case of duplicate section anchors as added bywikitext_headings_to_anchors()replace characters matched by
substitute_rewith one ASCII white spacestrip leading and trailing
strip_charslowercase
Note
This normalization is not perfect: it’s a trade-off between several ones, some of which may prevent from converging to a lowest common denominator. However, only extreme edge cases might be affected.
- Parameters:
- Return type:
Column- Returns:
the column of normalized section headings
- section_topics.pipeline.parse_wikitext(keep_lists_and_tables, minimum_section_size)[source]¶
Currying function that passes args to the underlying
parse()PySpark user-defined function (UDF).See also this gist.
- section_topics.pipeline.extract_sections(articles, keep_lists_and_tables, minimum_section_size)[source]¶
Extract section data from articles
- Parameters:
articles (
DataFrame) – a dataframe of article pages as output bylook_up_qids()keep_lists_and_tables (
bool) – whether to keep sections with at least one standard wikitext list or tableminimum_section_size (
int) – minimum content character length for the section to be considered
- Return type:
DataFrame- Returns:
DF of sections with links (1 row per link), DF of sections with images (1 row per section)
- section_topics.pipeline.normalize_wikilinks(link_column)[source]¶
Lowercase the first character of wikilink target titles.
The link target is case-sensitive except for the first character.
- Parameters:
link_column (
str) – a dataframe column name of wikilinks- Return type:
Column- Returns:
the column of normalized wikilinks.
Nonevalues are kept.
- section_topics.pipeline.handle_media(sections, media_prefixes)[source]¶
Separate media links from other ones.
Detect media links via lowercased lookup of namespace prefixes.
- Parameters:
sections (
DataFrame) – a dataframe of sections and wikilinks as output byextract_sections()media_prefixes (
list) – a list of namespace prefixes for media pages
- Return type:
Tuple[DataFrame,DataFrame]- Returns:
the dataframe of media links and the remainder of the
sectionsdataframe
- section_topics.pipeline.clean_up_links(sections)[source]¶
Filter empty strings and add a column of normalized wikilinks.
- Parameters:
sections (
DataFrame) – a dataframe of sections and wikilinks as output byextract_sections()- Return type:
DataFrame- Returns:
the cleaned
sectionsdataframe
- section_topics.pipeline.resolve_redirects(sections, redirects)[source]¶
Follow section wikilinks redirects.
If a wikilink points to a redirect page, replace its title with the redirected one. Normalize redirects via
normalize_wikilinks().- Parameters:
sections (
DataFrame) – a dataframe of sections and cleaned wikilinks as output byclean_up_links()redirects (
DataFrame) – a dataframe of page titles and redirected page titles as output byload_redirects()
- Return type:
DataFrame- Returns:
the dataframe of redirected wikilinks. Both original and normalized ones are kept.
- section_topics.pipeline.gather_section_topics(sections, articles)[source]¶
Align section wikilinks to their Wikidata QIDs: the so-called section topics.
- Parameters:
sections (
DataFrame) – a dataframe of sections and normalized wikilinks as output byresolve_redirects(). Pass the output ofclean_up_links()to skip wikilinks pointing to redirect pages.articles (
DataFrame) – a dataframe of articles as output bylook_up_qids()
- Return type:
DataFrame- Returns:
the dataframe of sections and topics (as Wikidata QIDs). Original topic titles are kept.
- section_topics.pipeline.compute_relevance(topics, level='section')[source]¶
Compute either the section-level or the article-level relevance score for every section topic.
The section-level score is a standard term frequency-inverted document frequency (TF-IDF). The article-level score is a custom TF-IDF, where TF is across wikis and IDF is within one wiki.
Workflow:
filter null topic QIDs
compute TF numerator: occurrences of one topic QID in a section or page QID
compute TF denominator: occurrences of all topic QIDs in a section or page QID
compute TF: numerator / denominator
join with input on section or page QID and topic QID
compute IDF numerator: count of sections or page QIDs in a wiki
compute IDF denominator: count of sections or page QIDs where a topic QID occurs, in a wiki
compute IDF: log( numerator / denominator )
join with input on wiki and topic QID
compute TF-IDF: TF * IDF
- Parameters:
topics (
DataFrame) – a dataframe of section topics as output bygather_section_topics()level (
str) – (optional) at which level relevance is computed,sectionorarticle
- Return type:
DataFrame- Returns:
the input dataframe with the
tf_idfcolumn added
- section_topics.pipeline.compose_output(scored_topics, all_topics, snapshot, page_namespace=0)[source]¶
Fuse scored topics with null ones and build the output dataset.
- Parameters:
scored_topics (
DataFrame) – a dataframe of scored topics as output bycompute_relevance()all_topics (
DataFrame) – a dataframe of all topics as output bygather_section_topics()snapshot (
str) – a weekly snapshot to serve as the constant value for thesnapshotcolumn of the output dataframepage_namespace (
int) – (optional) a page namespace to serve as the constant value for thepage_namespacecolumn of the output dataframe
- Return type:
DataFrame- Returns:
the final output dataframe
- section_topics.pipeline.parse(wikitext, section_level=SECTION_LEVEL, section_zero_heading=SECTION_ZERO_HEADING)¶
Note
This is the core function responsible for the data heavy lifting. It’s implemented as a PySpark user-defined function (
pyspark.sql.functions.udf()). It must be called by theparse_wikitext()currying function, which seems the only way to pass an optional denylist of section headings.Parse a wikitext string into a list of dicts. Each dict represents a section within the given wikitext and has the following keys:
index- section number,0for the leading sectionheading- section heading,section_zero_heading’s value for the leading sectionlinks- list of section wikilinksimages- list of section image file names
A section heading is normalized through
wikitext_headings_to_anchors(). Clean wikilinks with mwparserfromhell’s strip_code(). This can lead to empty strings that should be filtered.