<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>sigh.dev - #web-scraping</title><description>Scott Cooper on TypeScript, React, open source, San Francisco, and the web. - Posts tagged with &quot;web-scraping&quot;</description><link>https://sigh.dev/</link><item><title>Meta’s Muse is fantastic for web scraping</title><link>https://sigh.dev/posts/metas-muse-is-fantastic-for-web-scraping/</link><guid isPermaLink="true">https://sigh.dev/posts/metas-muse-is-fantastic-for-web-scraping/</guid><description>All-you-can-scrape for $80 a month.</description><pubDate>Thu, 01 Oct 2026 07:00:00 GMT</pubDate><content:encoded>&lt;p&gt;I haven’t found much use for the personal assistant part of Muse. People on X (Twitter) keep using these assistants to order DoorDash or something. I order DoorDash like five times a year, so that’s not going to speed anything up for me. I also don’t get many emails, and managing my appointments is easy enough already.&lt;/p&gt;
&lt;p&gt;What I do need is a bunch of data from Reddit and YouTube and a boatload of cheap AI tokens. That’s where Muse comes in handy. I’m building &lt;a href=&quot;https://fullsets.fm&quot; rel=&quot;noreferrer noopener&quot; target=&quot;_blank&quot;&gt;fullsets.fm&lt;/a&gt;, a side project that aggregates professionally recorded concerts (festival livestreams, Tiny Desk, etc.), organizes them by artist, and extracts tracklists. I’m looking for complete performances with good audio and video.&lt;/p&gt;
&lt;p&gt;I’ve got a pipeline that uses Luna and Gemini 3.8 via Gemini Studio to cheaply determine whether a YouTube video has the quality and content I’m looking for. What I don’t have is a good way to find the videos in the first place. I can scan a single YouTube channel for qualifying videos, but I need a way to find all the weird little channels, like T-Mobile Europe, that have for some reason uploaded a few random concerts. There are high-quality concerts scattered across all these little channels, and that’s part of why I built fullsets.fm (a work in progress).&lt;/p&gt;
&lt;p&gt;Muse helps me find those channels and videos. I’ve also got it at the other end of the pipeline, helping with artist matching, reviewing what Luna and Gemini output, and making a final decision to approve or reject each video.&lt;/p&gt;
&lt;p&gt;I can tell Muse to scan Reddit and YouTube for 1,000 videos matching my criteria, or build a list of artists who have ever played at Coachella and then go through them one artist at a time, searching YouTube for matching concerts. It sits there and spins until it’s done, punishing those websites with a machine in the cloud that has seemingly endless bandwidth and a decent CPU.&lt;/p&gt;
&lt;p&gt;I paid $80 for a month of Muse, and I’ve had it use a few billion tokens a week on this work. A lot of the time, it also seems not to count token usage at all. Their model (Muse Spark 1.3) is pretty mediocre at writing, and I would never use it to code, but I’m finding it really useful for this Gemini Flash-style workload. They seem happy to give me effectively unlimited tokens and a managed VM with Chrome.&lt;/p&gt;
&lt;p&gt;I worry about what will happen when someone points Muse at my own websites. These agents don’t sleep, and they can pretend to be real users extremely well. If it’s not fast enough, you can spin up more processes.&lt;/p&gt;
&lt;p&gt;Anyway, I see why &lt;a href=&quot;https://www.theverge.com/tech/998078/amazon-blocks-meta-muse-ai-agent-shopping&quot; rel=&quot;noreferrer noopener&quot; target=&quot;_blank&quot;&gt;Amazon blocked it already&lt;/a&gt;. Let’s say my goal was dropshipping or something. Could I scrape thousands of Amazon pages all day long? I don’t know where this leads, but we’ll likely see more websites block Muse unless Meta prevents abuse. There are still some early rough areas of Muse, jobs will freeze in the middle, communication with this somewhat shitty model can sometimes be frustrating and I’m not sure the UI makes much sense when running multiple jobs that are using the browser.&lt;/p&gt;
&lt;p&gt;If this post makes you want to try it out, my &lt;a href=&quot;https://muse.ai/join&quot; rel=&quot;noreferrer noopener&quot; target=&quot;_blank&quot;&gt;Muse&lt;/a&gt; referral code is &lt;code&gt;69WSDQ&lt;/code&gt;. Enter it in Settings within 48 hours of signing up and we both get a billion tokens.&lt;/p&gt;
&lt;figure data-breakout=&quot;wide&quot;&gt;&lt;a class=&quot;content-image-link&quot; href=&quot;./concert-refinement.png&quot; rel=&quot;noreferrer noopener&quot; target=&quot;_blank&quot;&gt;&lt;img alt=&quot;A retro computer classifying concert recordings in an empty office&quot; title=&quot;Click to open in new tab&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; width=&quot;1774&quot; height=&quot;887&quot; src=&quot;/_astro/concert-refinement.BHKU3j9T_ZE6FmC.webp&quot; srcset=&quot;/_astro/concert-refinement.BHKU3j9T_9dUYm.webp 640w, /_astro/concert-refinement.BHKU3j9T_2w0Wh3.webp 750w, /_astro/concert-refinement.BHKU3j9T_b94WY.webp 828w, /_astro/concert-refinement.BHKU3j9T_Z2i2eEN.webp 1080w, /_astro/concert-refinement.BHKU3j9T_2moAbb.webp 1280w, /_astro/concert-refinement.BHKU3j9T_Z2pvlA6.webp 1668w, /_astro/concert-refinement.BHKU3j9T_ZE6FmC.webp 1774w&quot; /&gt;&lt;/a&gt;&lt;figcaption class=&quot;image-subtext&quot;&gt;AI-generated image inspired by Severance&lt;/figcaption&gt;&lt;/figure&gt;</content:encoded><category>ai</category><category>web-scraping</category><category>fullsets</category></item></channel></rss>