AI Developmentr/LocalLLaMA

I classified 3.5M US patents with Nemotron 9B on a single RTX 5090 — then built a free search engine on top

Read original
ai-coding-toolslocal-llm-workflowsdomain-specific-aihybrid-search-architecture

Patent attorneys need exact phrase matching. 'solid-state battery electrolyte' should match those exact words, not semantically similar documents about 'energy storage.' FTS5 gives sub-second queries on 3.5M records with zero external dependencies.

Key takeaways

  • Local LLM (Nemotron 9B on RTX 5090) classified 3.5M patents in 48 hours, demonstrating consumer hardware can handle enterprise-scale classification tasks
  • Contrarian architecture choice: FTS5 full-text search outperforms vector embeddings for domain-specific use cases requiring exact phrase matching and deterministic results
  • Complete technical stack disclosed: SQLite FTS5 + local LLM query expansion + BM25 ranking with custom weights + FastAPI, hosted on Chromebook via Cloudflare Tunnel - proving production-grade search doesn't require cloud infrastructure
  • Patent lawyer with 1 month coding experience built production search engine, signaling democratization of AI tooling for domain experts without traditional engineering backgrounds
  • Hybrid approach: Uses LLM for natural language query expansion into boolean queries, then traditional search for retrieval - combining strengths of both paradigms

Why this matters for operators: Legal tech, enterprise search architecture, local LLM deployment for specialized domains, hybrid search strategies

I cover AI×GTM intelligence like this every Wednesday.

Get STEEPWORKS Weekly

More picks

AI DevelopmentLenny's Podcast

Humans will keep inventing new reasons why we must stay in the loop with agents

  • Human resistance to full AI autonomy is not purely technical—it's psychological and organizational; companies will rationalize keeping humans in decision loops even when agents are capable
  • The 'human-in-the-loop' requirement may become a self-perpetuating narrative rather than a genuine necessity, driven by organizational risk aversion and change resistance
  • Product leaders at scale (Notion) are observing this pattern, suggesting it's a widespread phenomenon across enterprise AI adoption, not isolated to specific use cases
ai-agent-adoptionhuman-in-the-loopai-governance
GTM Ops**RevOps Impact (Jeff Ignacio)

Comp plans for consumption pricing

  • Consumption pricing fundamentally breaks traditional SaaS comp models—requires rethinking sales incentive structures around usage vs. contract value
  • Four distinct contract structures exist (pay-as-you-go, uncommitted, committed, hybrid), each requiring different compensation mechanics and sales behaviors
  • Enterprise consumption-based deals create tension: customers want flexibility, sales teams need predictability for quota attainment—comp design must bridge this gap
revenue-platform-consolidationconsumption-pricing-modelssales-comp-design
AI×GTMGTM OS: The Future GTM Operator

3 revenue motions your AI is only half wired into

  • Model parity has arrived: OpenAI/Claude now trade evenly on core tasks, making 'better AI' a non-differentiator—the edge shifts to integration depth into existing revenue motions
  • Waste is quantified: teams paying $17K-$37K/month for AI seats that never touch pipeline generation; real cost is opportunity cost of unused capacity, not subscription fees
  • Lean teams have a structural advantage: cannot out-buy larger competitors on model access, but can out-embed them by wiring AI 1 revenue motion deep (pipeline → content → deals) with proprietary deal context competitors haven't seen
ai-sdr-adoptionrevenue-platform-consolidationback-to-basics-gtm

This analysis was produced using the STEEPWORKS system — the same agents, skills, and knowledge architecture available in the GrowthOS package.