ALIS: article-level image suggestions

This is ALIS (reads Alice): suggest images for Wikipedia articles that don’t have one.

Inputs come from Wikimedia Foundation’s Analytics Data Lake:

High-level steps:

Output pyspark.sql.Row example:

Row(
    page_id=9696852,
    id='ddd3bad8-327a-11ee-8991-f4e9d4472fd0',
    image='Gamez,_Cónsul_-_2018_Junior_Worlds_-_5.jpg',
    origin_wiki='commonswiki',
    confidence=96,
    found_on=['ruwiki'],
    kind=['istype-commons-category', 'istype-lead-image'],
    page_rev=132612416,
    section_heading=None,
    section_index=None,
    page_qid='Q21660678',
    snapshot='2023-07-24',
    wiki='itwiki',
)

More documentation lives in MediaWiki.

image_suggestions.alis.LIMIT_PER_QID = 20

Amount of suggestions per Wikidata QID

image_suggestions.alis.load_local_images(spark, short_snapshot)[source]

Load locally stored images through the image_suggestions.queries.local_images Data Lake query.

Parameters:
  • spark (SparkSession) – an active Spark session

  • short_snapshot (str) – a YYYY-MM date

Return type:

DataFrame

Returns:

the dataframe of wikis, file page IDs, and file names

image_suggestions.alis.load_suggestions_with_feedback(spark)[source]

Load image suggestions that were reviewed by users.

Parameters:

spark (SparkSession) – an active Spark session

Return type:

DataFrame

Returns:

the dataframe of wikis, page IDs, and image file names

image_suggestions.alis.get_illustratable_articles(spark, hive_db, snapshot)[source]

Collect Wikipedia articles that are suitable candidates for image suggestions.

A candidate has either no images or its images are used so widely across Wikimedia projects that they are probably icons or placeholders.

Parameters:
  • spark (SparkSession) – an active Spark session

  • hive_db (str) – a Data Lake’s Hive database name

  • snapshot (str) – a YYYY-MM-DD date

Return type:

DataFrame

Returns:

the dataframe of wikis, page IDs, page titles, and page QIDs

image_suggestions.alis.generate_suggestions(spark, hive_db, snapshot, limit_per_qid, coalesce)[source]

Gather the full ALIS dataset.

Parameters:
  • spark (SparkSession) – an active Spark session

  • hive_db (str) –

    a Data Lake’s Hive database name

  • snapshot (str) – a YYYY-MM-DD date

  • limit_per_qid (int) – an integer that limits the amount of suggestions per Wikidata QID. If > 0, suggestions are ordered by confidence. 0 stands for no limit

  • coalesce (int) – an integer to control the amount of files per output partition. A higher value implies more files but a faster and lighter execution

Return type:

DataFrame

Returns:

the ALIS dataframe