r/ProxyEngineering 19d ago

Discussion 💬 First BrightData and now NetNut? What's happening?

34 Upvotes

Ok so this is gonna be a bit of a ramble but I need y'all to hear this. I work adjacent to the scraping/data space (not naming employer, y'all know how it is) and I always kind of just you know, accepted that "residential proxies" were this slightly grey but mostly fine thing. Like yeah obviously somebody's home IP is being used, somebody agreed to something somewhere, moving on. Then yesterday I see netnut is just gone. Not down, not maintenance page, straight up FBI DOJ IRS-CI seizure banner, Google and Lumen and Shadowserver all stamped on it too like it's a whole coalition operation. Like bruh, I have never in my life seen this on any website. So imagine the shocker lol. And I'm sitting there like wait since when did the IRS care about proxy IPs??? Turns out IRS-CI does financial crime investigations and it could be related to that, but still, seeing that logo on a proxy provider's main website??? Diabolical mate.

So I go searching what's going on and apparently it's tied to this thing called the Popa botnet, basically it's been running for like four years hijacking Android TV boxes (remember that Bright Data SDK thingy?? And Krebs traced a chunk of it back to actual NetNut infrastructure through some ex-employee's domain. Not some randoms, an actual former VP of R&D there. Google's own blog post says they think the network was over 2 million devices at one point and that a lot of those "different" proxy brands you see around are just NetNut wearing a different logo through resellers. I swear I seen a post talking about how there is one or two proxy providers sharing those IPs/reselling whatever. Makes sense now, people?? Am I the only one concerned here with a few other nerds? And to be fair I want to be fair here, Alarum (the parent company) came out and started calling it that a botnet is inaccurate, also said that there's real KYC and consent flows and misuse detection on their end. So it's not like this is some minor thing everyone agrees on, it's actively disputed and I think that matters, I'm not trying to just slam a company because Krebs wrote a scary headline, no. Make what you will out of it, but I things are rising to the surface, more and more often. Again it does make me think about how much of the "residential proxy pool" any of us are touching day to day is actually consented in a way a normal person would recognize as consent, versus consented in the "buried in paragraph 14 of an SDK terms screen" way. Like where's the line between legit residential network and just a nicer branded botnet, and how would any of us even know the difference from the outside.

Not trying to start a witch hunt on the whole industry, genuinely just spent my evening reading legal filings and threat intel blogs instead of sleeping like a normal person, but after Bright Data SDK shennanigans and now Netnut, boy, there's gonna be more stuff coming out. So I figured I'd share since I know a bunch of you here actually work with this infra daily and probably have opinions on these matters.

Sources if anyone wants to go check themselves.

Krebs on Security, the original Popa botnet reporting: https://krebsonsecurity.com/2026/06/popa-botnet-linked-to-publicly-traded-israeli-firm/

Google's threat intel writeup on the disruption: https://cloud.google.com/blog/topics/threat-intelligence/google-continued-disruption-residential-proxy-networks

Alarum's official response: https://alarum.io/alarum-technologies-responds-to-inquiry-into-residential-proxy-networks/

And divinetworks.com itself, which is also showing the seizure page now: https://divinetworks.com/

P.S Some time later I found that they mixed up the domains because .com is down, but not .io which is the main page of netnut.

EDIT:

It would seem that their main website was also seized - https://netnut.io/


r/ProxyEngineering 19d ago

Hot Take 🔥 The NetNut FBI seizure raises a question nobody asks: where do residential IPs actually come from?

54 Upvotes

I went down the rabbit hole reading about the NetNut situation last night and honestly, the FBI seizure wasn't even the craziest part for me lol

The craziest part was realizing that somewhere out there, a guy is probably watching Netflix on his Samsung TV while his TV is simultaneously acting as infrastructure for somebody running a scraping operation on the other side of the world (I have a really old TV that and I don't watch it that much honestly, so I think I avoided this bullet)

LIKE WHAT? That's not even a joke anymore

According to the news coverage happening at the moment, researchers found proxy SDKs embedded in a huge number of smart TV applications. Not sketchy APKs from some random forum. actual apps running on LG and Samsung TVs

I think that in this case, the conversation quickly shifted from NetNut to a much bigger question that I don't think someo people don't think about:

Where do residential IPs actually come from???

Don't get me wrong, I don't think about this question either that much and it's a really complex answer to a complex question

The sourcing side always felt like somebody else's problem honestly

But after reading about the alleged connection between NetNut and the Popa botnet, I started realizing how little visibility most of us actually have into the supply chain behind residential proxies

What surprised me most wasn't the claim that millions of devices were involved. It's 2026. Every month there's another story involving millions of compromised devices, so unfortunately that part barely registers anymore

The surprise part for me happened when learning how many layers can exist between the person buying a residential proxy and the device providing that residential IP (like wth)

The part I keep coming back to is whether this is even a problem that can be solved completely

If residential proxy inventory passes through multiple layers of aggregators, partners, SDK providers and resellers, how many companies can honestly say they know the origin story of every IP in their network

At what point does a provider stop being a network operator and become a network consumer just like the rest of us?

What struck me while reading all this is that the proxy industry spends an enormous amount of time talking about performance metrics. We compare success rates, country coverage, session length, pool sizes, pricing, uptime and all the usual stuff. Yet I can barely remeber seeing a serious discussion about where these networks actually come from. Maybe that's because the answer is complicated, but after reading through the NetNut reporting it suddenly feels like one of the most important questions we could be asking.

SO very few people seem interested in tracing the supply chain behind the product itself, which is funny because that's ultimately the foundation everything else is built on

One thing that comes to mind in this situation is that some providers seem to be putting more effort into transparency than others and I've seen that happen pretty recently. I've seen providers like nodemaven, Iproyal or proxygonzo openly talk about IP quality filtering, network quality standards, and support processes. Others have started publishing more information about sourcing, partnerships, compliance policies, or how they acquire residential inventory in the first place. This should probably be the standard moving forward for everyone that's affected in this case

I'm not saying anyone deserves a free pass, and honestly this whole story makes me want to be more skeptical rather than less. But I do think there's a meaningful difference between providers that are willing to discuss where inventory comes from and providers that treat the entire supply chain as a black box.

Maybe that's where the industry needs to go next. We already compare success rates, uptime, sticky sessions and pool sizes. Maybe a few years from now we'll also be comparing transparency reports, sourcing disclosures, consent models and network provenance and being sceptical of each proxy provider?

Sources where I found this story: https://hivesecurity.gitlab.io/blog/netnut-popa-botnet-fbi-seizure-residential-proxy/

Also huge props to this user as he predicted the future lmao - https://www.reddit.com/r/ProxyEngineering/comments/1u00w9f/my_samsung_tv_is_literally_being_rented_out_as_a/


r/ProxyEngineering 15h ago

Hot Take 🔥 NetNut's parent company admitted they still don't know what happened and now a third of the staff might be gone

17 Upvotes

Since a lot of us watch this topic closely, here are some more news regarding Alarum Technologies. That's the company that owns NetNut, they put out another update on July 13 and it's kind of wild how little they're saying while saying a lot of words.

Here's a quick recap if you missed the earlier posts. Back on July 2 Alarum responded to press reports about its residential proxy network and paused some services as a precaution. Then on July 3 they disclosed that the FBI had actually seized domains tied to NetNut, and more kept getting seized as things went on. There was a July 4 update too. Now as of July 13, more than a week later, Alarum is saying they still don't know the root cause of the disruption. They've brought in an external cybersecurity and forensics team working under legal counsel, which tells you this isn't just an IT problem, this is a lawyered up problem.

This is where things start to get interesting. The operational efficiency plan they mention is reportedly going to affect close to a third of their workforce. Another interesting thing from an industry point is the framing. Alarum keeps saying they're investigating whether their own network was used for malicious or unlawful purposes by third parties, which puts them in the position of victim/cooperator rather than target, at least in their own telling. But domain seizures like this usually mean the government believes there's something to seize evidence from, and residential proxy networks have always had this legality question hanging over them about how the IPs are actually sourced. This might be the moment that forces the industry to answer that question directly.

What you guys think, is this a one-off enforcement action against a specific bad actor who abused NetNut's network, or is this the first step in a broader look at how residential proxy sourcing works across the web? If you've used NetNut recently, did you notice service disruptions before any of this became public?

Sources if you want to read the filings instead of just my summary:

https://www.globenewswire.com/news-release/2026/07/13/3326137/0/en/Alarum-Technologies-Provides-Further-Update-Regarding-Recent-Developments.html

https://www.globenewswire.com/news-release/2026/07/03/3321824/0/en/Alarum-Technologies-Provides-Update-Regarding-Recent-Law-Enforcement-Action.html

https://www.sec.gov/Archives/edgar/data/1725332/000121390026075179/ea029704201ex99-1.htm

Anyways, I'm still watching this case closely


r/ProxyEngineering 16h ago

Help 🆘 Need help finding a inexpensive static dedicated residential proxy provider

5 Upvotes

Hi, I’m trying to get started to use multiple accounts to advertise and have had a hard time finding a proxy provider that offers around 20 dedicated static residential proxies. I already have a anti-detect browser set up, so if you got any recommendations please let me know.


r/ProxyEngineering 12h ago

Hot Take 🔥 Too big to fail. What happened with NetNut

0 Upvotes

NetNut is facing a major crisis after researchers linked its proxy network to “Popa,” a botnet of more than two million compromised devices, including smart TVs and streaming boxes. In early July 2026, Google and the FBI began disrupting the infrastructure, while domains were reportedly seized.

Google said hundreds of threat groups routed traffic through suspected NetNut exit nodes. Reports suggest devices were enrolled through hidden SDKs in streaming, IPTV, and utility apps, often without clear consent.

The fallout includes millions of lost IPs, reduced reliability, account disruptions, reputational damage, and risks for resellers using NetNut’s infrastructure

How CyberYozh Ensures Proxy Quality and Security

CyberYozh protects proxy quality through ethical IP sourcing, continuous screening, and rapid abuse response. Its IP Checker detects fraud risks, blacklists, and suspicious network types before proxies are used. Residential and mobile IPs come from informed, consenting participants, while a dedicated abuse team investigates reports and removes questionable addresses quickly.


r/ProxyEngineering 1d ago

Guides What proxies do you need for scraping Instagram without getting flagged?

5 Upvotes

I buried myself deep inside Meta’s algorithms and figured out what actually triggers its safety systems.

The quick answer: Meta’s anti-bot stack scores every request based on IP reputation, request velocity, session consistency, and behavioral fingerprint.

So, what works better for safer scraping?

My honest answer is residential or mobile proxies.

Instagram’s anti-bot systems can flag datacenter IPs almost instantly because their ASN patterns and IP ranges are easy to identify. I usually start with rotating residential proxies for high-volume scraping, such as profiles, hashtags, and public posts. I escalate to mobile proxies only for endpoints where residential IPs start getting rate-limited or challenged. I also keep the request velocity around 20–40 requests per minute per IP, avoid aggressive concurrency, rotate sessions carefully, and make sure the behavioral fingerprint does not look fully automated.

That is the exact setup and rule set I personally follow when working with Instagram.

You are very welcome to share your own observations about Meta’s anti-bot stack.


r/ProxyEngineering 2d ago

Help 🆘 Best resis for Akamai

10 Upvotes

Anyone have a good provider for Akamai security?


r/ProxyEngineering 2d ago

Help 🆘 Proxy source

6 Upvotes

I am looking for an unlimited residential routing proxy source to supply my bot.


r/ProxyEngineering 2d ago

Help 🆘 Best Proxy for creating TikTok account in UK (company is located in Malaysia)

6 Upvotes

Hi everyone, I am a newbie to proxy and know nothing much about proxy, only VPN.

We are trying to create tiktok account that are based in UK region instead of Malaysia to get organic UK traffic.

What are the best proxy to use? Especially for mobile, since we would only use a mobile to create a TikTok account.


r/ProxyEngineering 2d ago

Discussion 💬 Do your proxy logs separate IPv4 and IPv6 exits?

5 Upvotes

I realized my benchmark logs record the exit address but never split results by IP version. That feels like a blind spot now that more networks and clients prefer IPv6 when it is available.

Older public sites, geo databases, and routing paths may not treat IPv4 and IPv6 identically. I do not have a result to claim yet, so the next test is simple: the same public URL set through IPv4-only and mixed routes, with Decodo, Byteful, and one dedicated datacenter range in the comparison. I will track parsed output, timeouts, latency, and geo agreement separately.

Has anyone already measured this on residential pools? I am curious whether IP version changes the result or merely exposes weaknesses in the client stack.


r/ProxyEngineering 2d ago

Guides How to build a deep research agent

6 Upvotes

Spoiler: it's not the model. It's not RAG. It's a markdown file.

Been some time since I shared about the builds so here it is. I built a deep research agent that takes open-ended questions, runs 15-30 tool calls over 10-25 minutes, and produces structured reports with sources. The model choice mattered less than I expected because the structure around it is where it's at.

Multi-model flow beats single-models. LangChain's open_deep_research repo (12k stars) does this well. They use gpt-4.1-mini for summarizing search results, gpt-4.1 for research reasoning, and a separate model for compressing accumulated findings before the context window blows up. This separation alone outperforms a single frontier model at everything. Their benchmarks on Deep Research Bench (100 PhD-level tasks, 22 fields): GPT-5 scores 0.49, default gpt-4.1 config scores 0.43 at ~$46 total, Claude Sonnet 4 hits 0.44 at $187. Do the math.

External state management is the savior. LLMs lose coherence after ~5 tool calls in a long-running sequence. Simplest implementation looks like this: the agent writes a markdown todo file at the start, updates it after each step, and re-reads it before deciding what to do next. This gives the model explicit visibility into what's done, what's in progress, and what's remaining. Without it, agents routinely skip subtasks or repeat work. Neat addition? I know :D Then LangGraph solves this at the structure level with built-in state machines and checkpointing. The todo file approach is more transparent and debuggable. LangGraph's approach seems to scale better.

My suggestion - don't dump full search results into context. Here's how I do it. Search → summarize to 2-3 sentences → store summary → proceed further. Pull full content only during synthesis. This dropped my token usage ~50% and improved output quality because the model wasn't being overwhelmed in irrelevant text. (Again, the token usage, not sure how accurately the LLMs give this data, but from my calculations it was around 50% In general, a typical deep research session consumes 50-150K input tokens and 10-30K output tokens.

Another thing to mind, LangSmith tracing. With it I could see the agent calling the same search query three times consecutively. You cannot debug agentic systems from outputs alone. You need the full tool implementation sequence, reasoning chain, and token counts per step.

My current setup is this: Claude Sonnet 4 for reasoning, gpt-4.1-mini for summarization, Tavily for search, you can also use Oxylabs Fast Search API for faster response times and you get structured results quickly. Choose what;'s better for you. Filesystem backend persisting intermediate findings as JSON. Runs as a LangGraph state machine. $2-4 per report.

Hope you'll find this interesting and useful.


r/ProxyEngineering 3d ago

Help 🆘 I am looking for an Proxy Testing Expert

4 Upvotes

Hi,

I am looking for a proxy testing expert who can do really test proxies for different purposes and share the data.


r/ProxyEngineering 4d ago

Discussion 💬 Maintaining 200+ scrapers is becoming unsustainable

7 Upvotes

I've been building scrapers for years. But mainly focus was on the single-targets. I'd choose a website, reverse engineer it then build something out of it and move on. That part I have no issues and I fell quite confident in it.

About 2 months or so I started a monitoring project (freelancing) that tracks product data across roughly 200 e-commerce sites. Getting the data was not a problem as mentioned, its where I am feeling confident, however, keeping everything running smoothly without any hiccups is a problem. For example, websites change their layouts constantly. A class name gets renamed. A div gets put one level deeper. A price field moves from the HTML to a JS-rendered component. (this is the worst), now since I monitor 200 sites and something breaks, is NOT SUSTAINABLE at all. It's just too much work to keep an eye on everything. Sure the project pays well, but do I want to stress over something like this everyday??

Right now I have CSS and XPath selectors stored per site in a config file. Usually, when something breaks I manually inspect and update the selector. Now with all those websites, it's crazy work. At this scale that's just not sustainable as I mentioned. I tried going back to LLM-based extraction when selectors fail but it only works maybe 50% of the time. Hallucinated prices are worse than no prices. I also built a simple alerting system that marks when the output schema doesn't match expected fields. It catches the issue but in reality it doesn't fix it. I have to do it myself.

I spend more time maintaining parsing part than I ever spent building the scrapers.

Please tell me that there is an easier way for this.

P.S Yes I am aware of dedicated scrapers and parsers but it's much more expensive than doing everything on my own. Plus I like the building part of the process. (Not the maintaining one). Has anyone found a good solution or at least some sort of an alternative?


r/ProxyEngineering 5d ago

Discussion 💬 What evidence actually gets a proxy refund approved?

4 Upvotes

Refund policies usually say some version of “case by case,” which is not very helpful once a project is already failing and everyone is annoyed. I’ve started keeping a small dayone baseline for every provider account. Same public URLs, timestamps, connection errors, returned content, and geo results byteful included. If performance changes later, I can show the difference instead of arguing from memory or comparing my logs with a dashboard that uses different success criteria. It also forces me to test the account while the purchase is still fresh, when a problem is easier to raise.
Has anyone received an actual refund from clean before-and-after logs, or do providers usually offer account credit instead?


r/ProxyEngineering 5d ago

Help 🆘 Bypassing anti-bot measures

3 Upvotes

This post is written with new users in mind. Write your own suggestions for the most optimal/best ways to bypass modern anti-bot solutions such as Cloudflare/DataDome/Kasada. What worked for you? What was your experience tackling these anti-bot measures, what was successful, what wasn't. Share your ideas. The mic is yours.


r/ProxyEngineering 5d ago

Reviews Testing CyberYozh Scraper UI vs broken DOMs. LLM self-healing, Playwright, and MCP.

4 Upvotes

I picked a few massive e-commerce targets with highly dynamic layouts. You know the drill. You write a clean Playwright script. The target updates its DOM structure overnight. Your entire pipeline crashes. You wake up to empty databases.

I ran these through the new CyberYozh Open Scraper update. I wanted to see how it handles this exact failure cycle. Spoiler alert. It changes the workflow entirely. You do not just write CSS selectors anymore. You tell AI agents what you need. They fix broken parsers mid-flight. But throwing raw LLMs at a crawler introduces new bottlenecks. Here is how the system actually handles them.

The Architecture & UI

Raw JSON testing is awful. The new update ships a Node.js interface on port 7000. Playwright handles browser rendering on the backend via port 8000. You drop in target URLs. You configure rendering timeouts. You align your network location using built-in proxy integration. No custom routing logic required.

The UI streams live crawl statistics via Server-Sent Events (SSE). You watch the target site map build itself as an indented tree. Because the system tracks everything in real-time, you see exactly what the crawler is doing. A single click stops in-flight requests gracefully.

LLM Self-Healing Parsing

Hardcoded CSS selectors represent a massive single point of failure. Targets mutate their HTML daily. This is where the engine actually works.

You execute your standard deterministic parser first. If a layout changes and a required field returns empty, the extraction does not abort. The AI automatically takes over.

```bash
curl -X POST http://localhost:8000/api/v1/scrape/preset/page \
  -H 'Content-Type: application/json' \
  -d '{
"source": "amazon_product",
"preset_params": {"asin": "B08N5WRWNW"},
"llm": {"model": "openai/gpt-4o-mini"}
  }'
```

Sending full e-commerce HTML to an LLM eats tokens instantly. It also crashes context windows. The engine strips scripts and structural boilerplate. For LLM processing, you request the `fit_markdown` format. The parser drops navigation menus. It ignores ads. It delivers pure content formatted specifically for model ingestion.

The model reads the cleaned data. It understands your required output schema. It locates the missing data points despite the broken selectors. But it does not just return the text. The system infers the new working CSS selector. It caches this selector for future runs. Your next request uses the standard deterministic parser again. You save massive API costs.

LLMs also hallucinate. The scraper forces strict JSON schema validation. If the model invents a price or missing stock status, the validation catches it instantly.

Server-Managed Authenticated Sessions

Scraping behind logins usually means dropping your context repeatedly. Open Scraper handles this with server-managed authenticated sessions. You authenticate once using a declarative JSON script.

```json
{
  "steps": [
{"op": "goto", "url": "https://example.com/login"},
{"op": "fill", "selector": "#username", "value": "$creds_email"},
{"op": "fill", "selector": "#password", "value": "$creds_password"},
{"op": "click", "selector": "button[type=submit]"}
  ]
}
```

The server maintains your persistent browser context. It tracks cookies securely. But proxies alone do not guarantee stable connections. The system actively manages browser fingerprints to match your aligned network location. You continuously scrape protected internal profiles without triggering re-authentication blocks.

If you hit aggressive device-verification challenges, you use the escape hatch. You inject pre-authenticated session cookies directly into the API to restore stable access.

Native MCP Integration

Writing custom middleware to bridge AI agents and scraping APIs wastes time. Both the scraper and crawler services eliminate this. They mount native Model Context Protocol (MCP) endpoints via Streamable HTTP.

You open the dedicated MCP tab in the UI. You configure the JSON arguments. You verify the output visually. Then you add the local server URL to your Claude Desktop config. You prompt the agent to crawl a target and summarize pricing pages. The agent submits the job, polls the server status, fetches the results, and parses the data autonomously.

Conclusions

Visual testing: Node.js UI handles complex setups locally.
Broken pipelines: LLM fallback fixes broken parsers mid-job and caches new selectors.
Cost control: Pre-processing trims heavy HTML before model inference.
Data integrity: Schema validation blocks model hallucinations.
Login walls: Server-managed sessions and fingerprinting keep your browser context active.
AI workflows: Native MCP endpoints connect directly to Claude and LangChain.

Has anyone else tested the self-healing features in production? I am curious about the token cost at scale once caching is fully optimized.


r/ProxyEngineering 6d ago

Hot Take 🔥 Bright Data locked down Residential access for new zones, self serve is basically dead now

20 Upvotes

Just saw this in their docs. It's over fellas. If you try to set up a new Residential zone after July 7, 2026, you now need to pass a human reviewed KYC check before you get access. Well done, guys!! No more instant setup, no more just adding a zone and going. You have to sign up under a registered company, verify a corporate email, and wait on their compliance team to approve you. Personal email accounts, Gmail, Outlook, whatever, are flat out not eligible anymore.

Existing zones are exempted in. If you already had a Residential zone running on or before July 7, you keep it and nothing changes on your end. This is only for the new zones going forward, and it applies across the board too, shared rotating IPv4, the IPv6 Mega Pool, dedicated Residential, all of it.

Feels like the natural next step after the Samsung/ LG TV SDK mess a few weeks back. They're clearly trying to get ahead of the scrutiny on how those residential IPs are sourced by tightening who can route traffic through them. Makes sense in theory, opted in home IPs are a liability if literally anyone with a card can rent them, but it's going to be a pain for smaller devs and freelancers who don't have a registered business entity sitting around.

Anyone here gone through this KYC process recently with them?


r/ProxyEngineering 6d ago

Discussion 💬 TIL your proxy provider has opinions about which websites you're allowed to scrape

5 Upvotes

Spent two days debugging a public court records scraper this week. Rotation fine, headers fine, target barely defended. Requests still dying on exactly one domain. Turns out the provider silently blocks .gov ranges in their acceptable use filter. No error saying so, the request just hangs. Great.

So I started checking this across providers I keep around. Ran the same court URL through Bright Data, Byteful and a smaller reseller and got three different behaviors. One connects, one 403s at the gateway, one times out with zero explanation. All three sell "unrestricted" residential plans.

I get why blocklists exist. But public records are public, the government publishes this stuff for people to read. Anyway, domain-level routing checks are now step one of every provider trial I run, before speed, before geo.

Does anyone keep a running list of who blocks what?


r/ProxyEngineering 6d ago

Guides 401 vs 407, quick guide to reading auth errors

Thumbnail
3 Upvotes

r/ProxyEngineering 6d ago

Discussion 💬 Have you ever asked for a refund from a proxy provider? And did you get it?

4 Upvotes

I was talking to a few people recently about proxy providers and something interesting came up which are refunds.

It seems like a lot of people only think about it after something goes wrong. You test a provider, things don’t perform the way you expected, and then you’re stuck wondering if you can even get your money back or are complaining to get your money back

From what I’ve seen, policies vary a lot. Some providers are pretty strict (no refunds once traffic is used), others are more flexible depending on the situation and sometimes can actually give you some traffic back like a quality guarantee, and sometimes supporrt dissappears when you bring it up lol

Made me curious how common this is.


r/ProxyEngineering 6d ago

Help 🆘 MCP server for web search

10 Upvotes

I've been testing web search MCP servers with Claude Code and Codex. None of them do what I actually need. Especially when debugging.

What I've tried:

Built-in web search tools - faults on JS-rendered sites. If your docs use React or Next.js, you get nothing useful and that's a waste of time.

Playwright MCP - While it works it's okay, but one search ate 10% of my context window. Not sustainable for multi-lookup sessions.

Exa.ai, Brave Search, Context7 - Pay-per-query or hard limits on free tiers. Not adding another subscription so my AI can read a webpage. Tbh getting pretty tired of these subscriptions on everything.

docs-mcp-server - Requires a persistent background server. I don't like that it needs a persistency. What I want is something that boots with npx and dies when the session ends.

To sum up: free or self-hostable, something that handles JS-rendered pages, returns clean structured content instead of raw HTML full of navigation bars, stays context-efficient, and works with Claude Code / Codex / Cursor without special setup.


r/ProxyEngineering 6d ago

Reviews Found this useful, hopefully you will too

Thumbnail
4 Upvotes

r/ProxyEngineering 7d ago

Hot Take 🔥 Two dedicated IPs from the same order can behave like completely different products

10 Upvotes

I used to treat a batch of dedicated datacenter IPs as interchangeable. Same provider, same order, same monthly price. One range kept producing clean public product rows while another range struggled on the same sites.

The difference was not speed. The weaker IPs shared a subnet with addresses that had a rough public reputation history. That made the whole order feel like a lottery, even though the dashboard showed identical products.

Now I test every new range separately before adding it to a job. I compare small batches from Bright Data, Byteful, and another provider, then keep the ranges that produce clean output on the real target. I also ask for a subnet swap before writing off the provider entirely.

Does anyone have a simple pre-purchase check for dedicated ranges, or is a small live test still the only honest answer?


r/ProxyEngineering 8d ago

Guides Your Proxy Works in cURL, So Why Does It Completely Fail in Your Browser?

Thumbnail
8 Upvotes

r/ProxyEngineering 8d ago

Discussion 💬 Fast Search API vs traditional Web Scraper APIs

17 Upvotes

Your experience, which ones are worth taking a look and trying out, which are not?

Which one would you recommend for personal use, for company wide tasks.

All opinions are welcome