The files they were accessing were similar in character to the files retrieved by the wiki agents, using similar tricks (r.jina.ai)
【方法】攻击者使用相似的技术方法(如r.jina.ai)访问文件,这表明AI代理可能遵循某种可预测的行为模式。识别这些模式可以帮助开发更有效的防御措施,但同时也表明AI代理可能被训练来执行特定类型的网络活动。
The files they were accessing were similar in character to the files retrieved by the wiki agents, using similar tricks (r.jina.ai)
【方法】攻击者使用相似的技术方法(如r.jina.ai)访问文件,这表明AI代理可能遵循某种可预测的行为模式。识别这些模式可以帮助开发更有效的防御措施,但同时也表明AI代理可能被训练来执行特定类型的网络活动。
In May & June 2025, Duke University Libraries (DUL) staff successfully implemented Anubis, a configurable open source web application firewall (WAF), in order to stave off persistent onslaughts of AI-related bot scraping activity. During this pilot period (May 1 - June 10, 2025), aggressive bot scraping led to extended outages for three critical library platforms (Duke Digital Repository, Archives & Manuscripts, and the Books & Media Catalog), and in each case, implementing Anubis mitigated the problem.
https://hdl.handle.net/10161/32990
Aery, Sean (2025). Anubis Pilot Project Report - June 2025. Retrieved from https://hdl.handle.net/10161/32990.
Opinion and Order. OCLC Online Computer Library Center, Inc. v. Anna's Archive (2:24-cv-00144). District Court, S.D. Ohio.
The Court is sympathetic to OCLC's situation: a band of copyright scofflaws cloned WorldCat's hard-earned data, gave it away for free, and then ignored OCLC when it sued them in this Court. But mindful that bad facts sometimes make bad law, the Court requests that an Ohio court intervene before this Court makes any new state tort, contract, property, or criminal law.
The Court resolves to CERTIFY the novel Ohio-law issues identified above to the Supreme Court of Ohio. Plaintiff's counsel and Matienzo's counsel are ORDERED to propose an order containing all the information Ohio Supreme Court Practice Rule 9. 02 requires by April 11, 2025. The parties may file their proposed orders separately, or, if they so choose, they may file one joint proposed order. The Court will finalize a certification order afterward.
OCLC's motion for default judgment is DENIED without prejudice. See Lammert v. Auto-Owners (Mut. ) Ins., 286 F. Supp. 3d 919, 928-29 (M. D. Tenn. 2017) (adopting this same disposition). Because the answers to the certified questions may also determine Matienzo's motion to dismiss under Federal Rule of Civil Procedure 12(b)(6), ECF No. 21, the Court DENIES without prejudice that motion too. See id. The Court invites the parties to reraise their motions after the certification proceeding. See id.
The Court also grants OCLC leave to amend its Complaint to correct any of the above-identified pleading deficiencies.
By hacking WorldCat.org, scraping and harvesting OCLC’s valuable WorldCat
This is a matter of some debate—notably the recent LLM web scraping cases.
If you are going to crawl sites you better use Ferrum or Vessel because you crawl, not test.
https://forum.newsblur.com/t/is-apify-the-best-scraper-for-sites-without-rss/9179
RSS Scraper tools: - Apify https://apify.com/ - RSSHub: https://github.com/DIYgod/RSSHub - RSS Bridge: https://github.com/RSS-Bridge/rss-bridge - Five Filters: https://createfeed.fivefilters.org/ - AWS release notes feed: https://dyn.tedder.me/rss/aws-release-notes.xml - Far Side: https://dyn.tedder.me/rss/farside/daily.json
List of others here: https://tedder.me/generated_news_feeds/
Source for: https://apify.com/page-analyzer
Westrupp, E., Greenwood, C., Fuller-Tyszkiewicz, M., Berkowitz, T., Hagg, L., & Youssef, G. J. (2020). Text Mining of Reddit Posts: Using Latent Dirichlet Allocation to Identify Common Parenting Issues [Preprint]. PsyArXiv. https://doi.org/10.31234/osf.io/cw54u
Facebook began as a (horny) web scraping project, as did Google and all other search engines.
Facebook... errrr.
Scrapism
Trying to understand how to scrape data (damn I hate that phrase... it makes me thinnk of some kind of test for colon cancer or something). This pertains to #clubcovid.
Archiving service with an emphasis on scholarly publishing.
Archiving pages that block it.
"The problem: the automated web browsing tools they want to use (commonly called “web scrapers”) are prohibited by the targeted websites’ terms of service, and the CFAA has been interpreted by some courts as making violations of terms of service a crime."
Good news for anyone who uses the Internet as a source of information: A district court in Washington, D.C. has ruled that using automated tools to access publicly available information on the open web is not a computer crime
Pingback: Legality of Extracting Publicly Available User-Generated Content – PromptCloud Pingback: How to Scrape Facebook Posts for Free Content Ideas Pingback: Facebook data harvesting—what you need to know (From Phys.org) – Peter Schwartz
important readings
Google doesn’t use the facebook API to scrape facebook; they just scrape it.
really?
This is an extremely important case to remember. It has implications for all Fb users who want to own their past.
WAIL in Electron,
The author of the defunct ArchiveFacebook addon.
Need proof? In Linkedin v. Doe Defendants, Linkedin is suing between 1-100 people who anonymously scraped their website. And for what reasons are they suing those people? Let's see: Violation of the Computer Fraud and Abuse Act (CFAA). Violation of California Penal Code. Violation of the Digital Millennium Copyright Act (DMCA). Breach of contract. Trespass. Misappropriation.
Linkedin lawsuit -- terrifying
Turn websites into structured data.
First, we view technology evolution as a three-stage cyclical process of adoption, appropriation, and repossession. Users drive adoption. Users and providers alternatively drive appropriation and repossession, as users lead appropriation, while providers react when reclaiming the resulting innova-tions. Second, we identify three appropriation modes—baroquize, creolize, and canni-balize—that represent increasing degrees of power contestation by users. And third, we identify three repossession modes—co-opt, combine, and block—that represent increas-ingly antagonistic reactions by providers and mirror users’ appropriation strategies.
El documento como árbol es una convención fija inicial, para lograr cierto movimiento en el desarrollo de la plataforma y las dinámicas alrededor de la misma, pero dicha convención puede ser móvil después (como se indicaba en el primer texto sobre Grafoscopio). Textos rizomáticos o laberínticos como los presentados en la literatura latinoamericana (Cortazar, Borges) podrían ser construidos con Grafoscopio una vez la convención inicial se mueva. Esto implicaría pasar por las sucesivas fases e incluso "canibalizar" Grafoscopio al final, con la ventaja de que las tensiones entre proveedores y usuarios no son tan fuertes, pues son los usuarios los que se están proveyendo de tecnología a sí mismos y cambiándola por el camino. Los lugares de tensión ocurren cuando se manifiesta el caracter político de sus usos, por ejemplo haciendo web scrapping que viola los contenidos de los términos de uso de un sitio web (citar caso de Twitter).
We shouldn’t have to create open data by scraping websites. This information should be already available, easily accessed and provided in a machine-readable format from the original providers, be they city councils or transportation companies. However, until there’s another option, we’ll always have scraping.