National Cyber Warfare Foundation (NCWF)

Inside Aliens_eye: ML-blended username enumeration across 840+ platforms


0 user ratings
2026-10-07 21:23:52
milo
Red Team (CNA)
"Inside

Aliens_eye is an AI-assisted OSINT username scanner that checks 840+ platforms with a blended ML/heuristic detector, intended for authorized investigations and defensive footprinting research.








Toolarxhr007/Aliens_eye — AI-powered OSINT username scanner covering 840+ social platforms with ML-blended detection
CategoryOSINT / username enumeration (Python, MIT license)
Primary UseMapping an authorized subject's username footprint across social platforms for investigations, brand protection, and OSINT research
Safe UseLegitimate OSINT research, authorized investigations, and educational use; the README's own disclaimer restricts it to legal, ToS-compliant research on subjects you are authorized to assess
Telemetry NoteScans originate thousands of lightweight HTTP requests from one IP (unless proxied via --tor/--proxy), which site-side rate limiting and WAF rules can trivially flag as automated enumeration

Aliens_eye is the rare username-enumeration tool that treats detection accuracy as an engineering problem rather than assuming an HTTP 200 means an account exists. Built in Python and published under the MIT license with roughly 4.3k stars, it asynchronously probes 840+ platforms for a given username and classifies each response as Found, Maybe, or Not Found using a blend of a trained model and a heuristic scoring engine. The packaging is refreshingly mature: a PyPI package, Docker support, CI badges, and a documented stable library surface via aliens_eye.api.


The core differentiator is the detection pipeline. Every HTTP response is reduced to a 30-dimensional feature vector covering status-code buckets, username placement in the path, title, meta and canonical tags, error and profile keywords, DOM structure such as images, forms, and CSS class shapes, structured-data signals like og:type and JSON-LD Person, response timing, redirect counts, and per-site fingerprints learned from previous scans. Two judges then vote: a weighted heuristic engine and a pure-Python logistic-regression model that ships with the package so no sklearn runtime dependency is needed. The shipped model blends probabilities at 0.9 * ml + 0.1 * heuristic, with Found above 0.620 and Not Found below 0.360; a missing model file silently falls back to heuristics defined in core/detector.py.


What earns the project credibility is the honesty of its own evaluation. The README states plainly that the shipped model was fit on 1,455 samples across 286 platforms and scored on 142 held-out platforms with precision 0.65, recall 0.51, and a 9% false-positive rate, and that on unseen platforms no configuration was statistically distinguishable from a plain status-code check on F1. That candor is unusual in this genre and frames results correctly: Found and Maybe are leads to verify, not findings. The aliens_eye selfcheck subcommand surfaces precision, recall, F1, and false-positive rate per site, and aliens_eye corpus record plus selfcheck --corpus enables reproducible replay evaluation against a frozen response corpus.


Beyond raw scanning, the correlation layer is where investigative value concentrates. With --correlate, hits are clustered into likely-same-person groups using avatar hashes, bios, shared links, and display names — the avatar-image matching requires the [correlate] extra for Pillow. --recurse-depth N follows usernames linked from discovered bios and re-scans them, expanding the pivot graph, while --domains checks whether username.com and similar TLDs are registered and live. Results export to JSON, CSV, HTML, Markdown, PDF, and graph formats including GEXF, Mermaid, and Maltego CSV, which slots directly into established analyst workflows.


The operational surface is thoughtfully scoped. --site github,reddit and --exclude-site constrain targets, --no-nsfw skips adult platforms, and drop-in site maps are supported by placing a JSON file like { "site_name": "https://site/{}" } into sites.d/. Network routing supports --tor (requiring a local Tor daemon) or arbitrary proxies via --proxy socks5://.... A --watch 6h --notify mode re-scans on an interval and POSTs changes to a webhook, and --resume file.jsonl checkpoints long scans. The README defangs its example webhook as https://hooks.example/aliens, a small but telling attention to detail.


Architecturally, the codebase is organized under src/aliens_eye/ into core/ (scanner, detector, analyzer, http, exporter, fingerprints), ml/ (inference, training, dataset collection), utils/ (a rich-based console layer), and data/ (sites.json, the trained model, ground-truth sets). A textual-based interactive browser ships behind the [tui] extra, and — notably — aliens_eye serve exposes the scanner as an MCP server so LLM agents can drive it, an early signal of how OSINT tooling is converging with agentic workflows. A Playwright fallback behind the [browser] extra handles JavaScript-heavy pages that return empty DOMs to plain HTTP clients.


Retraining is a first-class path rather than an afterthought. After installing [train] for scikit-learn, aliens_eye train collect gathers a dataset from ground-truth accounts plus randomly generated non-existent usernames, train fit produces a model.json, selfcheck --split holdout --model model.json scores it on unseen platforms, and --model loads it at scan time. Ground truth is split site-disjoint between data/selfcheck.json and data/eval_holdout.json, and the README warns that scoring against the train split measures fit, not generalization. The aliens_eye label command adds an active-learning loop for hand-labeling uncertain hits.


Comparative evaluation is built in too: aliens_eye eval ablate scores detector configurations with bootstrap confidence intervals, and eval external runs Sherlock, Maigret, and WhatsMyName rule sets against the same stored responses for apples-to-apples baselines — you fetch those projects' data files yourself. This framing positions Aliens_eye less as a Sherlock clone and more as a research platform for the detection problem itself, which is arguably the more durable contribution.


Installation is a single pip install aliens-eye, with optional extras like pip install "aliens-eye[browser]" for the Playwright fallback. Usage stays non-weaponized and analysis-oriented in the documented examples: aliens_eye username --correlate --format gexf style invocations that produce reports rather than attack output. The aliens_eye.api module offers a stable async library surface returning plain dicts shaped like the JSON report, making it embeddable in larger authorized tooling.


From a defender's perspective, the telemetry profile matters: a full scan fires hundreds of near-simultaneous lightweight GET requests for patterned URLs from a single IP, which is trivially visible to rate limiters and WAF anomaly rules unless routed through --tor or --proxy. Privacy teams should also note the tool's watch mode and correlation features when assessing how easily an organization's personnel footprint can be mapped by third parties — useful input for social-engineering risk briefings and exposure-reduction guidance.


Within its lane — authorized OSINT, investigator workflows, exposure reviews of one's own staff or brand — Aliens_eye is one of the more technically serious entries in the username-enumeration space, chiefly because it publishes its own failure rates and ships the machinery to reproduce and improve them. Treat its outputs as ranked leads to manually confirm, respect platform terms of service, and keep it pointed at subjects you are authorized to research.



Official project repository for arxhr007/Aliens_eye.

Download Tool

Educational analysis for authorized security professionals. Use only in controlled, authorized environments.






Source: OffensiveSec
Source Link: https://www.offsecblog.com/2026/10/inside-alienseye-ml-blended-username.html


Comments
new comment
Nobody has commented yet. Will you be the first?
 
Forum
Red Team (CNA)



Copyright 2012 through 2026 - National Cyber Warfare Foundation - All rights reserved worldwide.