Email threat detection

Most phishing filters guess. PhishLoop shows its work.

A three-stage cascade — deterministic rules, a calibrated LightGBM classifier, and multimodal QR/OCR analysis — scores every message and explains exactly why, in under half a second.

POST /predict QUARANTINE
subject"Account verification required"
sendersecurity@paypa1‑support.com
triggered_rulesKNOWN_BAD_DOMAIN, CREDENTIAL_FORM_ACTION
reason_codesURL_BRAND_MISMATCH
risk_score
0.97
97.1%
Stage 2 recall on held-out phishing mail
91.4%
Stage 2 precision — low analyst noise
98.85%
Stage 3 QR-classifier accuracy, AUC 0.999
<5ms
Stage 1 rule-engine response time
01

How it decides

Every message moves through the cascade in order. Each stage only sees what the last one couldn't resolve — clear-cut mail exits fast, and only the ambiguous, high-risk cases pay for deeper analysis.

1

Deterministic rules

DMARC / SPF / DKIM validation, URL blocklists, Unicode homoglyph detection, and known-bad signature matching. Obviously malicious or clean mail is resolved immediately; anything ambiguous escalates.

DMARC/SPF/DKIM · URL blocklists · homoglyph detection · DNS entropy
< 5 ms
2

LightGBM classifier

142 engineered features spanning text n-grams, URL structure, sender graph, attachment metadata, and message context feed a calibrated gradient-boosted model, with a SHAP explainer attached to every score.

142-dim features · probability calibration · SHAP explainer
~ 100 ms
3

Multimodal analysis

HTML/DOM inspection, a purpose-built QR-code CNN classifier with automatic decode, and OCR on embedded images — catching image-only and QR-based phishing that a text model alone would miss.

HTML/DOM inspection · QR-code CNN + decode · OCR
~ 300 ms
Σ

Risk fusion

Stage 2 and Stage 3 scores combine in log-odds space into one calibrated final_risk_score, with a weighted-average fallback when the fusion model isn't trained.

 

02

What's underneath

Detection is the core, but the parts that make it trustworthy in production go beyond a single model.

Explainability

Every verdict ships with triggered rules, reason codes, and top SHAP-ranked features — analysts see why a message was flagged, not just a score.

QR & image phishing

A purpose-built CNN decodes and scores QR payloads embedded in emails and attachments — 98.85% accuracy, AUC 0.999.

Adversarial training

Trained and evaluated against a human + LLM-generated adversarial email corpus — brand impersonation, executive impersonation, and QR-phishing tactics included.

Awareness simulation

A closed-loop program sends employees adaptive-difficulty phishing emails and scores every response — report, ignore, click, or credential entry.

Feedback loops

SOC-verified labels and trusted-user reports feed back into the model, weighted by how much that source has earned your trust.

Deployment

A FastAPI backend exposes /predict, /health, and /metrics. Ship it with a single docker-compose up.


03

Turn every near-miss into training

PhishLoop doesn't stop at detection. A closed-loop awareness program sends each employee simulated phishing emails calibrated to their current skill level, and every response updates their profile.

LEVEL 123 — CURRENT45

Trusted status unlocks after 20+ simulations with a ≥85% true-positive rate and ≤5% false-positive rate — from then on, that person's reports carry more weight in the live model.

Report phishing correctlyFlags a real phishing email +2.0
Ignore / mark safe correctlyLeaves a legitimate email alone +1.0
False-positive reportReports a legitimate email as phishing −0.5
Clicks a phishing linkEngages with a malicious link −2.0
Enters credentialsFully compromised on a simulation −6.0

04

Learning that doesn't stop at deployment

Detections feed back into the system in production, and the decision layer is being pushed further in active research.

Live

SOC & trusted-user feedback

Two feedback endpoints close the loop between analysts, employees, and the model — each weighted by how much that source has earned your trust.

  • 01 SOC reviews a quarantined email → verified label submitted at full weight
  • 02 A trusted user reviews a challenged email → label submitted at 0.50–0.90 weight
  • 03 A SOC review always takes precedence over a user-submitted label
Research

Adaptive RL decision overlay

An experimental four-action reinforcement-learning layer — allow, quarantine, challenge user, escalate to SOC — trained on the same feedback signals. It's an active research direction, not yet part of the production decision path.

  • → Trained against adversarial, LLM-generated phishing email
  • → Target: ≥95% recall, ≤15% SOC-escalation rate

05

Measured, not marketed

Evaluated on held-out test splits for each stage's model.

Model Accuracy Recall Precision AUC
Stage 2 — LightGBM 93.9% 97.1% 91.4% 0.896
Stage 3 — QR classifier 98.85% 99% 99% 0.999
Stage 1 — rules< 5 ms
Stage 2 — LightGBM~ 100 ms
Stage 3 — multimodal~ 300 ms

See it on your own traffic.

We'll run the cascade against real samples from your environment and show exactly why each one was scored the way it was — no data leaves your infrastructure.

Try it out — hello@nusoft.co We reply from hello@nusoft.co