Understanding String Similarity Panels at ICANN

The Internet Corporation for Assigned Names and Numbers (ICANN) plays a central role in maintaining the stability and security of the domain name system, particularly through its oversight of the introduction of new top-level domains (TLDs). One of the critical, though often misunderstood, components of this process is the work of String Similarity Panels. These panels are specialized review bodies tasked with assessing whether proposed new TLDs are visually, phonetically, or conceptually too similar to existing ones or to each other, with the aim of preventing user confusion, technical collisions, and malicious exploitation. Understanding the function, criteria, and impact of these panels requires a deep dive into both linguistic principles and the technical standards that govern internet naming.

At the heart of the String Similarity Panel’s mandate is the notion that confusion between domain strings could undermine trust and stability online. If two TLDs—say, .hotel and .hotell—are too alike in appearance or sound, users might visit the wrong site, fall victim to phishing, or misinterpret the source of an online service. In the most sensitive cases, such as country names or high-profile generics, this confusion could have political or economic consequences. To mitigate these risks, ICANN empowers expert panels to evaluate proposed strings not only through algorithmic comparison but also through linguistic, visual, and cultural analysis.

The assessment begins with a visual similarity review. This involves evaluating how closely a proposed string resembles existing TLDs when rendered in common typefaces and scripts. The panel considers the length of the string, the shape of individual characters, and the probability that a user could mistake one for the other, particularly in small fonts or low-resolution displays. This process is complicated by the presence of homoglyphs—characters that appear similar or identical across different scripts. For example, the Latin “a” and the Cyrillic “а” (U+0430) may be visually indistinguishable in most fonts, making a Cyrillic-based TLD a potential vector for impersonation if it overlaps with a Latin-script equivalent.

Phonetic similarity is also a key dimension of review. The panel considers whether the string, when spoken aloud, sounds too much like an existing TLD or applicant string. This is especially relevant in multilingual or international contexts where domain names may be used in voice search, radio advertisements, or word-of-mouth communication. To assess phonetic similarity, the panel relies on a combination of linguistic expertise and phonetic transcription systems, such as the International Phonetic Alphabet (IPA). If two strings differ in spelling but are homophones in a major language—such as .site and .sight in English—this may trigger a contention set or rejection.

Conceptual similarity adds a third, more abstract layer. Here, the panel evaluates whether two strings represent the same idea or semantic category, which could also cause confusion even if the visual and phonetic forms are distinct. For example, .doctor and .physician could be considered conceptually similar despite having different letters and sounds. This analysis often draws upon lexical databases, translation references, and cultural knowledge, particularly in the case of internationalized domain names (IDNs) where translations of the same term appear in different scripts. The intent is to avoid allowing two parties to control essentially identical semantic real estate in the root zone.

When a proposed TLD is found to be too similar to an existing string or another applied-for string, it may be placed into a contention set. This means it cannot proceed to delegation until the conflict is resolved, either through mutual agreement, auction, or withdrawal. If the similarity is with an already delegated TLD, the new application may be denied outright. This mechanism ensures that the namespace remains uncluttered by near-duplicate identifiers and prevents bad-faith actors from exploiting typographic or linguistic ambiguities for phishing or trademark abuse.

The panels themselves are composed of subject matter experts selected for their proficiency in linguistics, typography, and international naming systems. These experts are independent of the applicants and ICANN staff, ensuring neutrality in judgment. They work in tandem with string similarity algorithms, which provide an initial computational assessment based on edit distance, visual rendering scores, and character set overlap. However, the final determination rests with human evaluators, who can account for nuances such as script-specific behavior, character context, and culturally specific connotations that algorithms alone cannot detect.

The application of string similarity rules becomes even more complex in the context of IDNs. Because these domain names can be rendered in scripts such as Arabic, Cyrillic, Chinese, and Devanagari, the number of potential confusable strings increases exponentially. The String Similarity Panel must not only understand the character sets and linguistic rules of these scripts but also how they interact when scripts are mixed or when Punycode representations are used. A proposed domain in Cyrillic that looks nearly identical to a Latin-script brand may be rejected, even if the underlying Unicode code points are entirely distinct.

Critics of the string similarity process often point to its subjective nature. While visual resemblance and phonetic overlap are objectively definable in some cases, in others, the threshold for “similar enough to be confusing” remains a matter of interpretation. This has led to appeals and disputes in prior rounds of new gTLD applications, especially when applicants believe that their string’s rejection was based on inconsistent or overly conservative judgments. To address this, ICANN has iteratively refined its guidelines, clarified similarity criteria, and introduced transparency measures around panel composition and methodology.

Despite these efforts, the stakes remain high. A successful gTLD application can translate into millions of dollars in registration revenue, brand prestige, and control over a linguistic or market niche. Conversely, a failed application due to similarity objections can result in sunk costs, lost strategic opportunities, and legal battles. As a result, many applicants now conduct pre-submission similarity analyses using both proprietary tools and third-party consultants with experience in IDN and gTLD policy.

In summary, String Similarity Panels at ICANN serve as guardians of clarity and security in the global domain name space. By applying detailed linguistic, visual, and phonetic analysis to every new string, they help prevent confusion, brand infringement, and user deception. Their work underscores the deeply interdisciplinary nature of domain name governance—where computer science meets linguistics, and where a few letters can carry global consequences. For any entity considering a new TLD, understanding the intricacies of string similarity review is not merely advisable—it is essential for navigating the intersection of language, policy, and internet infrastructure.

You said:

The Internet Corporation for Assigned Names and Numbers (ICANN) plays a central role in maintaining the stability and security of the domain name system, particularly through its oversight of the introduction of new top-level domains (TLDs). One of the critical, though often misunderstood, components of this process is the work of String Similarity Panels. These…

Leave a Reply

Your email address will not be published. Required fields are marked *