Demo 03 / Data extraction / APIs
Site to API
Paste a public web page. The server fetches it, parses the HTML and hands back its tables, repeated lists and metadata as clean JSON — at an endpoint you can call yourself.
How it works
A public server that fetches any URL a stranger types is a classic way into a private network. Most of this app is the guard.
-
01
Check the address
http and https only, ports 80 and 443 only, no
user:pass@. Odd IPv4 spellings —2130706433,0177.0.0.1,0x7f.1— are normalised by the URL parser before anything is compared, so they can't sneak past as "not an IP". -
02
Resolve, vet, pin
DNS is resolved on the server and every answer must be public unicast: loopback, RFC 1918, link-local and the
169.254.169.254metadata service, CGNAT, multicast, reserved, ULA and IPv4-mapped IPv6 are refused. The socket is then pinned to the vetted IP, so a second DNS answer can't swap it (rebinding). -
03
Fetch under limits
At most 3 redirects, each followed by hand and re-vetted from step 01. 8 seconds in total, 2 MB streamed then cut off, HTML content types only. The peer address is checked once more after connecting. Per-visitor rate limit, 4 fetches at once, 5-minute cache.
-
04
Parse like a browser
HTML5 parsing via cheerio/parse5. Tables are expanded through
rowspan/colspanso cells line up with their headers. Siblings sharing a tag + class signature, three or more, become record lists, ranked by content. No AI — the same page gives the same JSON.