PR mi/34501 reports the following:
...
$ gdb -q \
-ex 'set charset UTF-8' \
-ex 'interpreter-exec mi2 "-break-insert -f foo' \
-ex quit
&"�\235\214�\217 No symbol table is loaded. Use the \"file\" command.\n"
...
$
...
The output is a bit odd, but that gets better if we use
'set print sevenbit-strings on':
...
&"\342\235\214\357\270\217 No symbol table is loaded. Use the \"file\" command.\n"
...
The output we see there is the error emoji:
...
$ gdb
(gdb) b foo
❌️ No symbol table is loaded. Use the "file" command.
...
More specifically, two utf-8 encoded unicode characters:
- Cross Mark [1]: 0xE2 0x9D 0x8C
- Variation Selector-16 (VS16) [2]: 0xEF 0xB8 0x8F
Now the question: is GDB doing something wrong?
I think we probably should encode unicode characters in MI error strings as
octal escapes, independent of the sevenbit-strings setting. This patch does
not address this part.
Then there's the question whether we should emit emojis in MI error strings in
the first place [3]. In principle they're unicode characters encoded in UTF-8,
and we can expect other such unicode characters in translated error strings.
But, given that MI has can_emit_style_escape () == false, and already filters
out ANSI escape sequences, I think it's reasonable to also disable emojis.
As for implementation, I introduced a function emoji_allowed alongside
can_emit_style_escape, which defaults to the value of can_emit_style_escape.
Tested on x86_64-linux.
Approved-By: Tom Tromey <tom@tromey.com>
Bug: https://sourceware.org/bugzilla/show_bug.cgi?id=34501