Sources currently blocked from activation β Difficult to collect with standard HTML scrapers (firecrawl + sitemap); next-level techniques are required.
β Bucket C β blocked (3)
All are SPA / JS pagination / non-standard site structures. sitemap.xml does not contain actual article URLs.
| Source | URL | Cause | |------|-----|------| | LG AI Research | lgresearch.ai/blog | JS pagination, only 1 article in sitemap, no RSS | | NAVER LABS | naverlabs.com/storyList | storyList SPA, sitemap 404, no RSS | | SK DEVOCEAN | devocean.sk.com | Category landing only, all 6 sitemap URLs are navigation, no RSS |
β Resolved in this round
- kakao-brain β Confirmed as a duplicate domain of
kakao-techand deprecated. Expanded thekakao-techfeed withtech.kakao.com/posts/feed(472KB full archive). - upstage β Activated 110 articles in sitemap.xml mode
- perplexity, naver-clova, kakaostyle β Activated with firecrawl + JS rendering
- anthropicΓ3, mistral, samsung, xai β Activated with firecrawl + static HTML
Next step options (remaining 3)
- API endpoint exploration β Discover XHR JSON endpoints via DevTools Network analysis
- playwright-python standalone collector β headless Chromium + BeautifulSoup
- External feed services β Proxies such as RSSHub, Feed Creator