Irrelevant data detection

A set of utility functions that identify irrelevant images and unsuitable article or section candidates.

image_suggestions.unillustratable.STRIP_CHARS = '!"#$%&\' *+,-./:;<=>?@[\\]^_`{|}~'

ASCII punctuation characters to be stripped from section titles. Include the ASCII white space, don’t strip round brackets.

image_suggestions.unillustratable.SUBSTITUTE_PATTERN = '[\\s_]'

All kinds of white space to be substituted for the ASCII one; underscores turn into spaces as well.

image_suggestions.unillustratable.UNILLUSTRATABLE_P31 = ('Q577', 'Q29964144', 'Q3186692', 'Q3311614', 'Q14795564', 'Q101352', 'Q82799', 'Q21199', 'Q28920044', 'Q28920052', 'Q13406463', 'Q4167410', 'Q22808320', 'Q98645843', 'Q17099416', 'Q100775261')

If an article’s Wikidata item is an instance of one of the items in this list, then it’s not suitable for getting suggestions.

image_suggestions.unillustratable.PLACEHOLDER_IMAGE_SUBSTRINGS = ('flag', 'noantimage', 'no_free_image', 'image_manquante', 'replace_this_image', 'disambig', 'regions', 'map', 'default', 'defaut', 'falta_imagem_', 'imageNA', 'noimage', 'noenzyimage')

Image file names containing these substrings are probably icons or placeholders.

image_suggestions.unillustratable.get_disallowed_substrings_regex(substrings=('flag', 'noantimage', 'no_free_image', 'image_manquante', 'replace_this_image', 'disambig', 'regions', 'map', 'default', 'defaut', 'falta_imagem_', 'imageNA', 'noimage', 'noenzyimage'))[source]

Build a regular expression to detect image file names that may be icons or placeholders.

Parameters:

substrings (tuple[str, ...]) – a tuple of substrings that indicate icons or placeholders in image file names

Return type:

str

Returns:

the regular expression that matches the given substrings

image_suggestions.unillustratable.get_allowed_suffixes_regex(suffixes=('.bmp', '.jpeg', '.jpg', '.png', '.tif', '.tiff'))[source]

Build a regular expression to detect image file extensions that typically hold valid images.

Parameters:

suffixes (tuple[str, ...]) – a tuple of suffixes that indicate valid image file extensions

Return type:

str

Returns:

the regular expression that matches the given suffixes

image_suggestions.unillustratable.read_section_images_parquet(spark, section_images_parquet)[source]

Load images available in all sections of all articles of all Wikipedias, as output by imagerec.article_images.

Parameters:
  • spark (SparkSession) – an active Spark session

  • section_images_parquet (str) – a HDFS path to a parquet generated by imagerec.article_images

Return type:

DataFrame

Returns:

the dataframe of:

  • item_id (string) - page Wikidata QID

  • page_id (string) - page ID

  • page_title (string) - page title, in original case and underscored

  • article_images (array<struct<heading:string,images:array<string>>>) - images per section per page

  • wiki_db (string) - wiki project

image_suggestions.unillustratable.get_section_images(spark, section_images_parquet)[source]

Explode a dataframe as loaded by read_section_images_parquet() for easier processing.

Parameters:
  • spark (SparkSession) – an active Spark session

  • section_images_parquet (str) – a HDFS path to a parquet generated by imagerec.article_images

Return type:

DataFrame

Returns:

the dataframe of:

  • wiki_db (string) - wiki project

  • page_id (string) - page ID

  • page_title (string) - page title, in original case and underscored

  • section_heading (string) - page section, in URL anchor format. More details in section_topics.pipeline.wikitext_headings_to_anchors()

  • image (string) - Commons image file name

image_suggestions.unillustratable.get_non_illustratable_item_ids(spark, hive_db, weekly_snapshot)[source]

Gather Wikidata QIDs that aren’t suitable for getting suggestions.

See UNILLUSTRATABLE_P31.

Parameters:
  • spark (SparkSession) – an active Spark session

  • hive_db (str) – a Data Lake’s Hive database name

  • weekly_snapshot (str) – a YYYY-MM-DD date

Return type:

DataFrame

Returns:

the dataframe of Wikidata QIDs

image_suggestions.unillustratable.read_denylist_parquet(spark, denylist_parquet)[source]

Load denylisted section titles.

Parameters:
  • spark (SparkSession) – an active Spark session

  • denylist_parquet (str) – a HDFS path to a parquet generated by section_topics.scripts.gather_section_titles_denylist

Return type:

DataFrame

Returns:

the dataframe of:

  • wiki_db (string) - wiki project

  • section_heading (string) - page section, in URL anchor format.

image_suggestions.unillustratable.get_non_illustratable_sections(spark, denylist_parquet, dataframe, wiki_column, heading_column)[source]

Gather all Wikipedia article section headings that aren’t suitable for getting suggestions.

Parameters:
  • spark (SparkSession) – an active Spark session

  • denylist_parquet (str) – a HDFS path to a parquet generated by section_topics.scripts.gather_section_titles_denylist

  • dataframe (DataFrame) – a dataframe of irrelevant section headings

  • wiki_column (Column) – a dataframe’s column of wikis

  • heading_column (Column) – a dataframe’s column of section headings

Return type:

DataFrame

Returns:

the dataframe of:

image_suggestions.unillustratable.get_images_in_placeholder_categories(spark)[source]

Load images that belong to the placeholder Commons category.

Parameters:

spark (SparkSession) – an active Spark session

Return type:

DataFrame

Returns:

the dataframe of:

  • cl_from (bigint) - Commons page ID

  • cl_to (string) - Commons category page title, in original case and underscored

  • cl_type (string) - 'file'

  • page_title (string) - Commons page title, in original case and underscored

image_suggestions.unillustratable.normalize_heading_column(column, substitute_pattern='[\\\\s_]', strip_chars='!"#$%&\\' *+, -./:;<=>?@[\\\\]^_`{|}~')[source]

Same as section_topics.pipeline.normalize_heading_column().

Return type:

Column