wp-includes/utf8.php:109// Valid strings come through unchanged.
'test' === wp_scrub_utf8( 'test' );
// Invalid sequences of bytes are replaced.
$invalid = "the byte xC0 is never allowed in a UTF-8 string.";
"the byte \u{FFFD} is never allowed in a UTF-8 string." === wp_scrub_utf8( $invalid, true );
'the byte � is never allowed in a UTF-8 string.' === wp_scrub_utf8( $invalid, true );
// Maximal subparts are replaced individually.
'.�.' === wp_scrub_utf8( ".\xC0." ); // C0 is never valid.
'.�.' === wp_scrub_utf8( ".\xE2\x8C." ); // Missing A3 at end.
'.��.' === wp_scrub_utf8( ".\xE2\x8C\xE2\x8C." ); // Maximal subparts replaced separately.
'.��.' === wp_scrub_utf8( ".\xC1\xBF." ); // Overlong sequence.
'.���.' === wp_scrub_utf8( ".\xED\xA0\x80." ); // Surrogate half. Note! The Unicode Replacement Character is itself a Unicode character (U+FFFD).Once a span of invalid bytes has been replaced by one, it will not be possible to know whether the replacement character was originally intended to be there or if it is the result of scrubbing bytes. It is ideal to leave replacement for display only, but some contexts (e.g. generating XML or passing data into a large language model) require valid input strings.$textstringstring function wp_scrub_utf8( $text ) { /* * While it looks like setting the substitute character could fail, * the internal PHP code will never fail when provided a valid * code point as a number. In this case, there’s no need to check * its return value to see if it succeeded. */ $prev_replacement_character = mb_substitute_character(); mb_substitute_character( 0xFFFD ); $scrubbed = mb_scrub( $text, 'UTF-8' ); mb_substitute_character( $prev_replacement_character ); return $scrubbed; }Introduced in 6.9.0. Unchanged from 6.9.7 through 7.1.0.
Signature, return type and hooks compared across 3 parsed releases.
src/wp-includes/utf8.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it.