Tags: search*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Usenet-Rewind is a search platform that indexes Usenet newsgroup archives dating from 1981 to the present, covering early tech support discussions, scientific and academic exchanges, hobbyist communities, and firsthand accounts of internet history. Operated by Erie Data Systems, LLC, the service currently holds approximately 1.09 billion messages with a retention period exceeding 45 years.
    - 12.37 million new messages added recently
    - Offers API access for programmatic search
    - Supports advanced filtering by title, body, author, message ID, newsgroup, and date range
    - Provides a browsable timeline of Usenet topics, people, and events by decade
  2. Dan Russell writes about the power of AI-augmented search to retrieve hard-to-find information, using an example of finding a study on how the gender of lab assistants affects experimental outcomes on lab mice. He demonstrates how a simple query with AI can yield relevant results, leading to original source papers. The study highlights the impact of experimenter gender on reproducibility in scientific research.
  3. The Brave LLM Context API provides an advanced web search service specifically designed to ground Large Language Models (LLMs) in RAG pipelines or agentic workflows. It delivers pre-extracted content—such as text, tables, and code snippets—in a compact format optimized for machine consumption rather than human reading. Users can manage context through configurable token budgets and refine results using relevance thresholds or custom source ranking via Goggles.

    - Supports location-aware queries including point-of-interest (POI) and map data.
    - Features freshness filtering based on page modification or publication dates.
    - Includes a "strict" threshold mode to prioritize high-relevance content over breadth.
  4. A post-retrieval temporal layer designed to improve RAG systems by addressing time-blindness in vector searches. This library implements validity filtering, document kind classification, and exponential decay scoring to ensure retrieved information is fresh and accurate. It functions downstream of existing vector search systems without requiring re-indexing or new infrastructure.
  5. >"How I added temporal awareness and freshness tracking to a RAG system with no sense of time."
    2026-05-11 Tags: , , , , by klotz
  6. Dan Russell shares an observation from a recent diving trip regarding a peculiar behavior where two different fish species swim in tight formation, such as a Spanish hogfish being closely followed by a Trumpetfish.

    The post poses research questions to the community about the name of this phenomenon, its biological purpose, and which other combinations of species might exhibit similar patterns.
  7. Tavily is a powerful API connecting AI agents to the live web for real-time search, extraction, research, and web crawling. It provides a production-grade retrieval stack to ground LLMs with fresh, factual web context, reducing hallucinations.

    Built for scale, Tavily handles millions of requests with low latency and built-in safeguards against PII leakage and prompt injection. Trusted by over one million developers and major enterprises like MongoDB and IBM, it offers seamless integration with leading LLM providers for sophisticated AI applications.
    2026-04-10 Tags: , , , , by klotz
  8. Google has introduced Google-Agent, a new entity appearing in server logs, to differentiate between traditional search crawling (like Googlebot) and AI-driven content fetching triggered by user interactions. Unlike Googlebot which proactively crawls and indexes the web, Google-Agent operates reactively, only fetching content in direct response to user prompts within Google AI products. A key distinction is that Google-Agent ignores `robots.txt` directives, behaving more like a standard web browser due to its user-initiated nature. This shift necessitates that developers adapt their infrastructure to identify and manage Google-Agent traffic correctly, focusing on real-time request management rather than traditional crawl budgets.
  9. This article discusses how to conduct long-term research effectively using AI as a partner, moving beyond single-prompt queries. It emphasizes the need for "Long-Term Triangulation" – a continuous, iterative methodology. The author outlines four key pillars: building a persistent memory for the AI, tracking shifts in the AI's understanding, actively critiquing its responses with contradictory data, and performing meta-audits to identify blind spots in the research process. The goal is to foster productive friction and avoid intellectual echo chambers, ensuring both the human and the AI think critically.
  10. discrawl mirrors Discord guild data into a local SQLite database, allowing you to search, inspect, and query server history independently of Discord. It’s a bot-token crawler – no user-token hacks – and keeps your data local. It discovers accessible guilds, syncs channels, threads, members, and message history, maintains FTS5 search indexes for fast text search (including small attachments), records mentions, and tails Gateway events for live updates with repair syncs. It provides read-only SQL access for analysis and supports multi-guild schemas with a simple single-guild default. Search defaults to all guilds, while sync and tail default to a configured default guild or fan out to all discovered guilds if none is set.
    2026-03-08 Tags: , , , , , , , , by klotz

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "search"

About - Propulsed by SemanticScuttle