Millions of Phishing Emails Hid Their Most Suspicious Words in Plain Sight
The recipient saw a familiar word such as “funding.” The email filter could see fun, an unusual invisible character and ding.
Microsoft has documented a high-volume phishing campaign that inserted Unicode tag characters inside financial lure terms. The technique altered the underlying text without changing how the word appeared to a person reading the message.
The character range became prominent through research into ASCII smuggling and prompt injection against AI systems. The campaign shows the technique crossing back into conventional phishing at scale. Microsoft recorded between one million and 2.37 million messages on weekdays during the most intense period.
One invisible character changed the token stream
The campaign used characters from the Unicode Tags block, U+E0000 through U+E007F. In Microsoft’s example, an invisible U+E0020 TAG SPACE appeared inside the word “funding.” It is another example of the parser disagreement that also made a trusted software mirror become the phishing site: systems and people were making different trust decisions about the same content.
This was not a hidden instruction for an AI assistant. The attackers were using a single invisible separator to fracture terms that signatures, regular expressions or machine-learning tokenisers might otherwise recognise.
A secure pipeline can normalise or remove the character and recover the visible word. A literal detector may fail because the expected sequence of letters is no longer contiguous. A tokeniser may split the word into unfamiliar fragments and reduce the signal available to a classifier.
The trick borrowed attention from AI security
Invisible characters are not new to spam. Attackers have long used zero-width spaces, soft hyphens, non-breaking spaces and look-alike letters to defeat naive matching.
What changed was the character block and the scale. Unicode tag characters were comparatively obscure until researchers used them to demonstrate hidden prompt injection. That attention made tooling and examples easier to find. The phishing campaign applied the same parser disagreement to a much older objective: getting a financial lure past mail controls.
This is a recurring security pattern. A technique does not stay confined to the product category where it first attracts attention. Once public research makes a primitive understandable and reproducible, attackers can test it against every system that interprets the same data differently.
Legitimate marketing infrastructure complicated the picture
The messages used disposable finance-themed domains but were relayed through infrastructure associated with ActiveCampaign, a legitimate email-marketing platform. Links were rewritten through the platform’s shared tracking domains.
That does not make the platform malicious. ActiveCampaign told Microsoft that its moderation system gives messages containing invisible Unicode characters the same verdict as unobfuscated equivalents and treats heavy use as suspicious. The broader lesson is that abuse of reputable shared infrastructure can weaken simple reputation-based decisions.
It is the same operational problem seen with JWR’s live phishing sessions: defenders need to follow the complete behaviour, not treat one familiar service or apparently normal rendering as proof of safety.
Microsoft says about 98.5 percent of measured messages matched the campaign’s envelope pattern, while roughly 99.8 percent matched that pattern or the platform’s tracking-URL structure. Those are useful clustering signals, but shared sending ranges and tracking domains should not be blocked as if every tenant were hostile.
The attack also created a high-confidence defensive signal
Unicode tag characters are rare in ordinary mail. Their presence can therefore be more suspicious than the financial word the attacker tried to conceal. The main benign exception highlighted by Microsoft is regional flag emoji composition, which defenders can account for.
Microsoft provides Advanced Hunting queries for Defender customers, including searches for the tag-character range in message subjects and bodies. The company also publishes pivots based on sender naming patterns, envelope structure and tracking URLs.
What email-security teams should test
- Verify that mail-processing, sandboxing, DLP and security-classification pipelines remove or consistently normalise Unicode tag characters before signature and model evaluation.
- Test the exact content that downstream systems receive. A clean result at the first gateway does not prove that journaling, ticketing, archiving or AI-assistant integrations interpret the same string identically.
- Hunt for U+E0000 through U+E007F in inbound subject and body content, excluding legitimate flag-emoji cases.
- Use finance-themed disposable domains, shared-platform envelope patterns and tracking URLs as corroborating evidence, not as standalone block indicators.
- Render suspicious messages safely and compare visible text with source code and normalised text.
- Teach analysts that “what the user sees” and “what the model tokenises” can be different security objects.
The campaign peaked months before Microsoft’s publication and later declined sharply. That does not make the technique historical. The research gives defenders a reusable test for phishing, content moderation and any workflow where one parser hands text to another.
Sources and further reading
- Microsoft Security: ASCII smuggling crosses over from AI prompt injection to phishing evasion, published 3 September 2026. No publication time was provided.


