Irrelevant data detection¶
A set of utility functions that identify irrelevant images and unsuitable article or section candidates.
- image_suggestions.unillustratable.STRIP_CHARS = '!"#$%&\' *+,-./:;<=>?@[\\]^_`{|}~'¶
ASCII punctuation characters to be stripped from section titles. Include the ASCII white space, don’t strip round brackets.
- image_suggestions.unillustratable.SUBSTITUTE_PATTERN = '[\\s_]'¶
All kinds of white space to be substituted for the ASCII one; underscores turn into spaces as well.
- image_suggestions.unillustratable.UNILLUSTRATABLE_P31 = ('Q577', 'Q29964144', 'Q3186692', 'Q3311614', 'Q14795564', 'Q101352', 'Q82799', 'Q21199', 'Q28920044', 'Q28920052', 'Q13406463', 'Q4167410', 'Q22808320', 'Q98645843', 'Q17099416', 'Q100775261')¶
If an article’s Wikidata item is an instance of one of the items in this list, then it’s not suitable for getting suggestions.
- image_suggestions.unillustratable.PLACEHOLDER_IMAGE_SUBSTRINGS = ('flag', 'noantimage', 'no_free_image', 'image_manquante', 'replace_this_image', 'disambig', 'regions', 'map', 'default', 'defaut', 'falta_imagem_', 'imageNA', 'noimage', 'noenzyimage')¶
Image file names containing these substrings are probably icons or placeholders.
- image_suggestions.unillustratable.get_disallowed_substrings_regex(substrings=('flag', 'noantimage', 'no_free_image', 'image_manquante', 'replace_this_image', 'disambig', 'regions', 'map', 'default', 'defaut', 'falta_imagem_', 'imageNA', 'noimage', 'noenzyimage'))[source]¶
Build a regular expression to detect image file names that may be icons or placeholders.
- image_suggestions.unillustratable.get_allowed_suffixes_regex(suffixes=('.bmp', '.jpeg', '.jpg', '.png', '.tif', '.tiff'))[source]¶
Build a regular expression to detect image file extensions that typically hold valid images.
- image_suggestions.unillustratable.read_section_images_parquet(spark, section_images_parquet)[source]¶
Load images available in all sections of all articles of all Wikipedias, as output by
imagerec.article_images.- Parameters:
spark (
SparkSession) – an active Spark sessionsection_images_parquet (
str) – a HDFS path to a parquet generated byimagerec.article_images
- Return type:
DataFrame- Returns:
the dataframe of:
item_id (string) - page Wikidata QID
page_id (string) - page ID
page_title (string) - page title, in original case and underscored
article_images (array<struct<heading:string,images:array<string>>>) - images per section per page
wiki_db (string) - wiki project
- image_suggestions.unillustratable.get_section_images(spark, section_images_parquet)[source]¶
Explode a dataframe as loaded by
read_section_images_parquet()for easier processing.- Parameters:
spark (
SparkSession) – an active Spark sessionsection_images_parquet (
str) – a HDFS path to a parquet generated byimagerec.article_images
- Return type:
DataFrame- Returns:
the dataframe of:
wiki_db (string) - wiki project
page_id (string) - page ID
page_title (string) - page title, in original case and underscored
section_heading (string) - page section, in URL anchor format. More details in
section_topics.pipeline.wikitext_headings_to_anchors()image (string) - Commons image file name
- image_suggestions.unillustratable.get_non_illustratable_item_ids(spark, hive_db, weekly_snapshot)[source]¶
Gather Wikidata QIDs that aren’t suitable for getting suggestions.
See
UNILLUSTRATABLE_P31.
- image_suggestions.unillustratable.read_denylist_parquet(spark, denylist_parquet)[source]¶
Load denylisted section titles.
- Parameters:
spark (
SparkSession) – an active Spark sessiondenylist_parquet (
str) – a HDFS path to a parquet generated bysection_topics.scripts.gather_section_titles_denylist
- Return type:
DataFrame- Returns:
the dataframe of:
wiki_db (string) - wiki project
section_heading (string) - page section, in URL anchor format.
- image_suggestions.unillustratable.get_non_illustratable_sections(spark, denylist_parquet, dataframe, wiki_column, heading_column)[source]¶
Gather all Wikipedia article section headings that aren’t suitable for getting suggestions.
- Parameters:
spark (
SparkSession) – an active Spark sessiondenylist_parquet (
str) – a HDFS path to a parquet generated bysection_topics.scripts.gather_section_titles_denylistdataframe (
DataFrame) – a dataframe of irrelevant section headingswiki_column (
Column) – adataframe’s column of wikisheading_column (
Column) – adataframe’s column of section headings
- Return type:
DataFrame- Returns:
the dataframe of:
wiki_db (string) - wiki project
section_heading (string) - page section, in URL anchor format. More details in
section_topics.pipeline.wikitext_headings_to_anchors()
- image_suggestions.unillustratable.get_images_in_placeholder_categories(spark)[source]¶
Load images that belong to the placeholder Commons category.
- Parameters:
spark (
SparkSession) – an active Spark session- Return type:
DataFrame- Returns:
the dataframe of:
cl_from (bigint) - Commons page ID
cl_to (string) - Commons category page title, in original case and underscored
cl_type (string) -
'file'page_title (string) - Commons page title, in original case and underscored
- image_suggestions.unillustratable.normalize_heading_column(column, substitute_pattern='[\\\\s_]', strip_chars='!"#$%&\\' *+, -./:;<=>?@[\\\\]^_`{|}~')[source]¶
Same as
section_topics.pipeline.normalize_heading_column().- Return type:
Column