Adjudicating String Similarity Trends and Case Studies in the 2026 gTLD Program
- by Staff
The 2026 new gTLD program reintroduces a critical and often contentious aspect of domain name adjudication: the evaluation of string similarity. At its core, string similarity assessment determines whether two or more applied-for gTLD strings are so visually, phonetically, or semantically alike that their coexistence in the DNS would cause confusion for users or risk undermining the stability and trustworthiness of the domain name system. ICANN’s commitment to mitigating user confusion and preserving namespace clarity has led to the reimplementation—and refinement—of the String Similarity Review process that played a significant role during the 2012 round. As applicants once again engage with this process in 2026, historical precedents and emerging trends reveal both the evolving sophistication of the review mechanism and the complex interplay of linguistic, legal, and business considerations that inform outcomes.
ICANN’s String Similarity Review in 2026 retains the fundamental architecture from previous rounds. Applications undergo an initial assessment by an independent String Similarity Panel, which evaluates each proposed TLD against existing delegated gTLDs, reserved names, and other applications in the same round. Strings are considered confusingly similar if the average internet user would likely perceive them as indistinguishable or easily confusable. This includes visual similarity in Latin or non-Latin scripts, phonetic likeness in spoken form, and conceptual overlap where the meanings or intended use cases of strings are substantially alike. Strings found to be too similar are placed in a contention set, from which only one may ultimately be delegated unless withdrawn or resolved through contention resolution processes such as auctions.
Case studies from the 2012 round continue to inform the approach and expectations of applicants in 2026. One of the most instructive examples involved the .hotels versus .hoteis applications. Despite being in different languages—English and Portuguese respectively—the String Similarity Panel found that the visual and phonetic proximity between the two terms was sufficient to place them in a contention set. This decision triggered criticism from some linguistic experts and highlighted the tension between linguistic diversity and global user confusion. In response, ICANN’s 2026 round incorporates a more nuanced multilingual evaluation protocol, leveraging expanded linguistic panels to analyze similarity within script families, taking into account contextual usage, local pronunciations, and script directionality in right-to-left languages such as Arabic and Hebrew.
Another key precedent influencing the 2026 adjudications is the .inc versus .ink contention. Though ultimately resolved through private resolution, the initial grouping of the strings into a contention set showcased the potential for confusion arising not just from literal character sequence, but from typographic and semantic closeness. This example emphasizes that string similarity assessments cannot rely solely on Levenshtein distance or simple algorithmic comparisons; human judgment remains essential, particularly in cases involving short acronyms, culturally significant terms, or variants in stylization.
In the current round, additional scrutiny has emerged around Internationalized Domain Names (IDNs). The expansion of support for IDNs in the 2026 gTLD round, including increased demand for Arabic, Cyrillic, Devanagari, Chinese, and Hangul script strings, requires adjudicators to understand local typographical conventions and linguistic equivalencies. For example, a Cyrillic script gTLD that visually resembles an existing Latin script TLD—such as a Cyrillic .сом being confused with the Latin .com—will almost certainly be rejected or forced into contention, based on the established principle that even cross-script visual similarity can cause material user confusion. To address this, ICANN has expanded its use of script-based panels with regional linguistic expertise and developed new automated tools that simulate browser rendering across different character sets to test for visual mimicry.
The introduction of homograph protection measures in the 2026 program further influences how string similarity is adjudicated. With growing awareness of phishing risks and homoglyph attacks, applications are now subject to analysis against known homograph variants in both existing and proposed TLDs. This includes not only common Latin script swaps—such as using “rn” for “m” or “l” for “I”—but also more sophisticated cross-script confusables. Applications flagged by these checks may be required to demonstrate mitigations such as usage restrictions, branding differentiation, or community safeguards that reduce the likelihood of abuse.
Semantic similarity is another increasingly prominent factor. While less deterministic than visual or phonetic tests, semantic overlap can raise significant concerns when two TLDs aim to serve the same community or subject matter. For instance, a proposed .eco and .green may be found to be semantically similar enough to warrant contention, particularly if both claim environmental stewardship as their mission. In such cases, the String Similarity Panel considers not just the string itself, but also the application’s stated use case, branding plans, and target audience. Applicants are encouraged to provide detailed rationales, community letters of support, and evidence of distinct market positioning to avoid or mitigate findings of semantic confusion.
Importantly, the 2026 program also introduces new appeal and review mechanisms. Applicants subject to a string similarity finding may request a reevaluation through the Independent Review Process (IRP) or escalate procedural issues to the ICANN Board Accountability Mechanisms Committee. This provides applicants with more tools to contest decisions they view as inconsistent or procedurally flawed. To be successful, however, appellants must offer compelling evidence that the panel’s decision was unreasonable, lacked expert input, or failed to apply ICANN’s guidelines correctly. These mechanisms add procedural transparency, but also underscore the importance of strategic application planning from the outset.
To navigate string similarity risk in 2026, many applicants are engaging linguistic experts, legal advisors, and digital branding consultants early in the application drafting process. Pre-application assessments can identify potential conflict strings, simulate panel evaluations, and develop fallback branding plans should a string be forced into contention or rejected. Applicants with similar strings often explore joint ventures or negotiated withdrawals to avoid costly contention resolution, including ICANN auctions of last resort. In other cases, applicants may pivot to alternative strings that retain brand or mission alignment but avoid similarity pitfalls.
In summary, adjudicating string similarity in the 2026 gTLD program remains a complex, multilayered process balancing algorithmic assessments with human linguistic and contextual interpretation. The process has matured since 2012, incorporating global linguistic expertise, cross-script safeguards, semantic analysis, and more robust appeals mechanisms. Applicants must recognize that success in this arena requires more than creativity in string selection—it demands rigorous due diligence, proactive risk mitigation, and a deep understanding of how strings function across languages, scripts, and user expectations. As the global namespace continues to diversify, the principle of clarity in the public interest remains paramount, and the adjudication of string similarity will continue to shape the landscape of safe, intuitive digital navigation.
You said:
The 2026 new gTLD program reintroduces a critical and often contentious aspect of domain name adjudication: the evaluation of string similarity. At its core, string similarity assessment determines whether two or more applied-for gTLD strings are so visually, phonetically, or semantically alike that their coexistence in the DNS would cause confusion for users or risk…