Keenable SELECT is pitching a different division of labor for AI web research: let an agent express the job in SQL, push search, filtering and extraction into the query engine, and return structured rows rather than a pile of links.
The system is presented as an MCP server with a primary `select` tool. It accepts a read-only DuckDB `SELECT` statement, executes web and semantic operations outside DuckDB, feeds their outputs back into a row set, and runs the final SQL query. Its showcase includes research reports alongside the full query and tool trajectory used to produce them.
What changes
A conventional search workflow gives an AI agent a ranked list of pages. The agent then has to open, read and synthesize those pages—work that can consume substantial context and model tokens.
Keenable’s approach treats web retrieval as a table source. A query can search the web, fetch pages, filter rows with conventional SQL conditions, use semantic tests to identify relevant passages, extract named fields, normalize values and group the results.

For example, a request to identify researchers who moved between frontier AI labs could use `WEB_SEARCH` to retrieve pages, `SEM_MATCH` to retain pages describing a qualifying move, and `SEM_EXTRACT` to pull fields such as researcher, former lab, new lab and move month.
The available operators include:
- `WEB_SEARCH`, which runs multiple searches, merges ranked results and deduplicates URLs.
- `WEB_FETCH`, which retrieves specified pages as Markdown.
- `SEM_MATCH`, an LLM-based predicate for meaning-based filtering.
- `SEM_EXTRACT` and `SEM_EXTRACT_ALL`, which return one or multiple values from a row.
- `SEM_SCORE`, an embedding-based ranking signal.
- `SEM_NORM`, intended to assign a common key to semantically equivalent values for grouping.
The company says conventional SQL filters run before the LLM-powered operators. That sequencing matters: a precise `WHERE` clause can reduce the number of rows that require semantic model calls.
Why operators and builders should care
The product’s core proposition is not that SQL replaces judgment. It is that many repeatable research steps—retrieval, deduplication, eligibility filters, field extraction and aggregation—can be made explicit and inspectable.
For teams building internal research agents, this could offer a more controllable pattern than asking a general-purpose model to browse freely and compose an answer from its reading. A query provides a compact representation of the research logic, while stored result sets enable later queries to build on prior work.
That may be particularly useful for workflows such as market mapping, competitor monitoring, supplier discovery, policy tracking or talent intelligence, where the desired output is a table with defined columns rather than a prose summary alone.
Keenable also separates data gathering from presentation. A research agent operates in a tool loop using `select`; a second report agent receives a brief and result-set data, then creates an HTML report in a sandboxed Python environment. The system renders drafts, returns screenshots and JavaScript error counts, and lets the report agent revise within a fixed budget before publishing a shareable link.
The practical caveats
Structured output does not automatically mean verified output. Web pages can be incomplete, stale or contradictory, and semantic extraction can miss details or infer a wrong field. Users will still need to inspect source material, define inclusion rules carefully and validate consequential findings.
Cost and latency are also likely to depend on query design. Keenable says a single call can search more than 1,000 pages and that exact filters can run without LLM cost, but semantic matching and extraction are invoked on rows that survive those filters. Broad retrieval paired with vague predicates could still create expensive or noisy research runs.
What to watch next
The meaningful test will be whether teams can turn common research requests into reliable query templates, with clear provenance and predictable operating costs. The published trajectories are a useful step toward auditability: they expose not just the final report, but the queries, tool results and result sets behind it.
If that transparency holds up in real business workflows, SQL-shaped research could become a practical interface for agents that need to produce reusable datasets—not merely polished answers.



