Search indices

We generate a dataset that enables queries against Wikimedia Foundation’s search indices. It serves two purposes:

  • inject Image sources into Commons

  • deliver all available image suggestions to Wikipedias

Commons

Build full and delta datasets of weighted tags for Commons’ search index.

Images can receive tags from 3 sources, as output by image_suggestions.wikidata_and_lead_images:

The full dataset is stored in the image_suggestions.shared.SEARCH_INDEX_FULL_TABLE, and the delta in the image_suggestions.shared.SEARCH_INDEX_DELTA_TABLE Hive table of Wikimedia Foundation’s Analytics Data Lake.

image_suggestions.commons.build_weighted_tags(wd_data, li_data)[source]

Build the full state of a Commons search index’s weighted tags dataset.

Parameters:
  • wd_data (DataFrame) – a dataframe of Wikidata claims as output by shared._load_wikidata()

  • li_data (DataFrame) – a dataframe of Wikipedia article lead images as output by shared.load_lead_images()

Return type:

DataFrame

Returns:

the dataframe of weighted tags

image_suggestions.commons.write_weighted_tags(spark, tags, hive_db, snapshot, coalesce, delta_threshold)[source]

Write the full and delta datasets to Hive.

Don’t write the delta if its row count is greater than a given threshold.

Parameters:
  • spark (SparkSession) – an active Spark session

  • tags (DataFrame) – a dataframe of weighted tags as returned by get_commonswiki_file_data()

  • hive_db (str) – an output Hive database name

  • snapshot (str) – a YYYY-MM-DD date

  • coalesce (int) – an integer to control the amount of files per output partition. A higher value implies more files but a faster and lighter execution

  • delta_threshold (int) – an integer row count threshold that determines whether to include the delta

Return type:

Tuple[DataFrame, DataFrame]

Returns:

the full and delta dataframes

Wikipedias

Build full and delta datasets of boolean flags for all Wikipedias’ search indices, indicating whether an article has an image suggestion.

Flags follow weighted tags’ syntax, namely recommendation.image/exists|1 and recommendation.image_section/exists|1 for ALIS: article-level image suggestions and SLIS: section-level image suggestions respectively.

The full dataset is stored in the image_suggestions.shared.SEARCH_INDEX_FULL_TABLE, and the delta in the image_suggestions.shared.SEARCH_INDEX_DELTA_TABLE Hive table of Wikimedia Foundation’s Analytics Data Lake.

image_suggestions.wiki_indices.load_suggestions(spark, hive_db, snapshot, target)[source]

Load image suggestions from image_suggestions.queries.SUGGESTIONS_TABLE, as output by image_suggestions.shared.save_suggestions().

Parameters:
  • spark (SparkSession) – an active Spark session

  • hive_db (str) – a Hive database name

  • snapshot (str) – a YYYY-MM-DD date

  • target (str) – whether to load ALIS or SLIS. Accepted values: {'alis', 'slis'}

Return type:

DataFrame

Returns:

the dataframe of image suggestions

image_suggestions.wiki_indices.build_exists_tags(target, suggestions)[source]

Build the dataset of boolean flags that feeds all Wikipedias’ search indices, in the form of exists|1 weighted tags.

Parameters:
  • target (str) – whether to build tags for ALIS or SLIS. Accepted values: {'alis', 'slis'}

  • suggestions (DataFrame) – a dataframe of suggestions as returned by load_suggestions()

Return type:

DataFrame

Returns:

the dataframe of boolean flags

Raises:

ValueError – if target isn’t accepted