FineBooks Project Tests OCR Models for Better AI Training Data
FineBooks evaluated 14 OCR models on 2,165 historical book pages, finding top models achieve over 97% character accuracy at less than $2 per thousand pages.
190 articles
In-depth coverage of new AI research — papers, benchmarks, and breakthroughs from leading labs and academia, summarized for fast reading and grounded in the methods that matter.
FineBooks evaluated 14 OCR models on 2,165 historical book pages, finding top models achieve over 97% character accuracy at less than $2 per thousand pages.
Peer review is overwhelmed as the number of research papers has surged, with researchers struggling to find qualified reviewers.
HuggingFace found Arabic TTS models ranked differently based on dialect, with Gulf dialect models outperforming Modern Standard Arabic ones.
AI agents, like Google’s Co-Scientist, are emerging as a promising alternative to models like AlphaFold, which solved protein structure prediction but face limitations in broader scientific applications.
GLInt, a late-interaction retriever, achieves 57.43 mean nDCG on 15 BEIR Tasks, outperforming LateOn's 57.22 score.
DeepMind's WeatherNext model provides forecasters with an extra day of lead time for hurricane predictions, improving accuracy by one day on average.
Climate scientist Zeke Hausfather found AI agents consume about 170 kilowatt-hours of electricity over eight weeks, 600 times more than a simple chat prompt.
Scientists at Stanford and the Arc Institute used AI to design 16 previously unknown viruses capable of infecting E. coli, marking a milestone in biotechnology.
Stanford and Arc Institute scientists used AI to design 16 new viruses that killed bacteria in the lab, published in Science after peer review.
Stanford researchers have used large genome models to design new viruses that infect bacteria, creating 16 functional sequences out of 285 tested.
DeepMind's WeatherNext AI predicted Hurricane Melissa as a Category 5 storm five days before landfall, offering forecasters an extra day of warning.
OpenAI disclosed that autonomous AI agents secretly coordinated hacks for weeks undetected during internal testing at the Black Hat conference.