25,240 questions
Score of -3
0 answers
85 views
How do I know if I will be able to display correctly a country flag on my console log with Unicode characters, from my Spring-Boot application? [closed]
Unicode has set countries flags from regional codes. With two Unicode characters, you make them appearing... "if your console log allows it".
i.e.
\u1F1EB\u1F1F7 should produce the country ...
Score of -3
2 answers
315 views
Unicode escapes in Java and compilation [closed]
I have a few questions about translations of Unicode escapes in Java :
As per JLS 25, section 3.3 :
A Java compiler should use the \uxxxx notation as an output format to display Unicode characters ...
Score of 5
1 answer
205 views
What endianness does validate_utf32() use?
In simdutf, for Unicode encodings that have endianness, there are normally separate functions for Little Endian and Big Endian variations, for example: validate_utf16le() and validate_utf16be(), but ...
Score of 6
1 answer
128 views
LENGTH reports three characters for the two-symbol engineering label 𝛥P
A test-rig gateway stores engineering display labels in Apache IoTDB 2.0.8 tree model whenever channel configuration changes. One differential-pressure channel arrives as 𝛥P: a mathematical italic ...
Score of 0
0 answers
84 views
How to set consistent font fallback for mixed English/Hindi text in PptxGenJS?
I'm generating a PowerPoint programmatically with PptxGenJS (Node.js) that mixes English and Hindi (Devanagari) text on the same slide — for example, a bilingual title like "Email Migration Guide ...
Score of 0
4 answers
419 views
How can I check whether a string contains only digits in Go? [duplicate]
In Python, I can use the str.isdigit() method to check whether a string contains only numeric characters.
For example:
"12345".isdigit() # True
"123a".isdigit() # False
&...
Score of 1
1 answer
150 views
_kbhit() unusual behavior with Unicode codepage on Windows console
I'm setting a Windows console with code page UTF-8 (65001), but the results with _kbhit() are a bit erratic (described below), when inputting keys in the "Latin Supplement" range, and even ...
Score of 0
3 answers
220 views
Best alternative for ❌ (U+274C; Cross Mark)
So I am trying to have a simple enough error icon and obviously Unicode has some crosses and stuff there but the issue is that the one that makes sense in context, is annoyingly colored and looks ...
Score of 0
1 answer
68 views
Is Apple's UCCompareCollationKeys() a strong or a weak ordering?
I'm wondering whether the result of comparing collation sort keys (= already pre-processed strings for faster collation-compatible sorting/searching) is a strong or weak ordering. The implementation ...
Score of 1
2 answers
231 views
Can font-size be set by unicode range?
I'd like to change the size of the symbols but they cannot be wrapped in additional HTML tags. Is there a way to target a unicode range when setting font-size? Such as,
span.name[...unicode range...] ...
Score of 1
1 answer
162 views
How to search and replace unicode characters with a Word macro?
I am trying to replace several Unicode characters in text strings in Microsoft Word.
The issue is when I try to use these text strings in other applications, the Unicode characters convert to a ...
Score of 0
3 answers
93 views
How to get the name of characters made up of multiple codepoints with ICU?
I managed to find u_charName() for getting the name of a single character, but what about characters like flag emojis, which are made up of multiple codepoints? Do characters like that even have names?...
Score of 0
2 answers
147 views
OCR output contains “garbage” characters after special symbols (mojibake / control chars) — how to reliably clean before returning from LLM?
I have an on-prem OCR pipeline that returns extracted text inside a JSON blob. I parse the LLM response and call a local normalizer before returning the text to callers. Example call site:
result = ...
Score of -4
1 answer
227 views
Getting weird results from java string codepoints on a windows machine [closed]
package edu.practice.zapper;
import java.io.IOException;
import java.lang.ProcessBuilder.Redirect;
import java.nio.charset.Charset;
import java.util.ArrayList;
import java.util.Base64;
import java....
Score of 0
1 answer
97 views
Using XSLT3.0 in Saxon-JS 2, how can one configure the processor so that it accepts codepoints-to-string(8)?
Within Saxon-J, I can set the processor configuration to allow XML1.1 characters, for example by:
processor.getUnderlyingConfiguration().setXMLVersion(XML11);
I'm looking for the equivalent in Saxon-...