function search_simplify

Simplifies a string according to indexing rules.

Parameters

$text: Text to simplify.

Return value

Simplified text.

See also

hook_search_preprocess()

7 calls to search_simplify()
SearchQuery::parseSearchExpression in drupal/core/modules/search/lib/Drupal/search/SearchQuery.php
Parses the search query into SQL conditions.
SearchSimplifyTest::testSearchSimplifyPunctuation in drupal/core/modules/search/lib/Drupal/search/Tests/SearchSimplifyTest.php
Tests that search_simplify() does the right thing with punctuation.
SearchSimplifyTest::testSearchSimplifyUnicode in drupal/core/modules/search/lib/Drupal/search/Tests/SearchSimplifyTest.php
Tests that all Unicode characters simplify correctly.
SearchTokenizerTest::testNoTokenizer in drupal/core/modules/search/lib/Drupal/search/Tests/SearchTokenizerTest.php
Verifies that strings of non-CJK characters are not tokenized.
SearchTokenizerTest::testTokenizer in drupal/core/modules/search/lib/Drupal/search/Tests/SearchTokenizerTest.php
Verifies that strings of CJK characters are tokenized.

... See full list

File

drupal/core/modules/search/search.module, line 415
Enables site-wide keyword searching.

Code

function search_simplify($text, $langcode = NULL) {

  // Decode entities to UTF-8
  $text = decode_entities($text);

  // Lowercase
  $text = drupal_strtolower($text);

  // Call an external processor for word handling.
  search_invoke_preprocess($text, $langcode);

  // Simple CJK handling
  if (config('search.settings')
    ->get('index.overlap_cjk')) {
    $text = preg_replace_callback('/[' . PREG_CLASS_CJK . ']+/u', 'search_expand_cjk', $text);
  }

  // To improve searching for numerical data such as dates, IP addresses
  // or version numbers, we consider a group of numerical characters
  // separated only by punctuation characters to be one piece.
  // This also means that searching for e.g. '20/03/1984' also returns
  // results with '20-03-1984' in them.
  // Readable regexp: ([number]+)[punctuation]+(?=[number])
  $text = preg_replace('/([' . PREG_CLASS_NUMBERS . ']+)[' . PREG_CLASS_PUNCTUATION . ']+(?=[' . PREG_CLASS_NUMBERS . '])/u', '\\1', $text);

  // Multiple dot and dash groups are word boundaries and replaced with space.
  // No need to use the unicode modifer here because 0-127 ASCII characters
  // can't match higher UTF-8 characters as the leftmost bit of those are 1.
  $text = preg_replace('/[.-]{2,}/', ' ', $text);

  // The dot, underscore and dash are simply removed. This allows meaningful
  // search behavior with acronyms and URLs. See unicode note directly above.
  $text = preg_replace('/[._-]+/', '', $text);

  // With the exception of the rules above, we consider all punctuation,
  // marks, spacers, etc, to be a word boundary.
  $text = preg_replace('/[' . Unicode::PREG_CLASS_WORD_BOUNDARY . ']+/u', ' ', $text);

  // Truncate everything to 50 characters.
  $words = explode(' ', $text);
  array_walk($words, '_search_index_truncate');
  $text = implode(' ', $words);
  return $text;
}