Market record — California

    AI Visibility in San Francisco

    In San Francisco and Silicon Valley, retrieval is dominated by technical and funding corpora, so an entity's machine-readable footprint — docs, repositories, structured records — often outweighs its marketing site.

    Retrieval context

    This market produces more machine-readable evidence per entity than any other: documentation, changelogs, code hosts, funding databases, and technical writing. Models trained and grounded on that material describe companies in the vocabulary of those sources. It also moves fast enough that stale records are a live risk — pivots, renames, and shutdowns propagate into answers slowly and incorrectly.

    Conditions that decide inclusion

    Documentation as authority
    Technical documentation is retrieved and quoted more readily than marketing copy; it is the highest-leverage surface in this market.
    Record staleness
    Funding and company databases lag reality, and models repeat their stale version of an entity long after it changes.
    Category crowding
    Dozens of entities claim identical category language, so undifferentiated positioning is functionally invisible to a retrieval system.

    Sectors competing for citation

    • Software and AI infrastructure
    • Venture capital and startups
    • Developer tooling
    • Enterprise SaaS
    • Biotech and research

    Questions this market asks answer engines

    • How do startups get mentioned in ChatGPT and Claude answers?

    • Does documentation improve AI visibility for software companies?

    • Why is my company's AI-generated description out of date?

    Scope note

    Jason Todd Wade does not operate an office, storefront, or local business in San Francisco and Silicon Valley, California. This page is a research record about how AI systems retrieve and describe entities in this market. Work is remote and market-agnostic.

    Related guides

    Fig. 03 — Sources

    Sources and notes

    Every source cited here is national or platform-level. No study, dataset, or vendor documentation measures answer-engine behavior for San Francisco specifically, and none is implied to: the retrieval mechanics are the same everywhere, while the competitive set and the questions asked differ. Local observations on this page are descriptions of the market's entity landscape, not measured rankings, and carry no claim of local presence.

    1. [01]

      AI features and your website — Google Search Central

      Platform documentation

      Google states that AI Overviews and AI Mode draw on its regular web index, that standard indexing eligibility governs inclusion, and that preview controls such as nosnippet and max-snippet apply to AI experiences.

    2. [02]

      Top ways to ensure your content performs well in Google's AI experiences on Search — Google Search Central Blog, 2025

      Platform documentation

      Google's own guidance for AI experiences: no separate AI ranking system to optimize for, unique and satisfying content, technical crawlability, and accurate structured data.

    3. [03]

      Introduction to structured data markup — Google Search Central

      Platform documentation

      Structured data must describe content visible on the page; Google documents JSON-LD as the recommended format and describes how markup is used to understand page content.

    4. [04]

      Person — Schema.org

      Specification

      The Person type and its sameAs property, the vocabulary used here to bind one canonical entity node to its off-site profiles.

    5. [05]

      Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., arXiv (NeurIPS 2020), 2020

      Research

      The paper that introduced retrieval-augmented generation — the architecture behind why retrieval eligibility, not ranking position, determines whether a source can appear in a generated answer.

    6. [06]

      OpenAI crawlers and user agents — OpenAI

      Platform documentation

      OpenAI documents distinct user agents — OAI-SearchBot for search surfacing, ChatGPT-User for user-triggered fetches, GPTBot for training — each controllable independently in robots.txt.

    Verify this yourself

    Machine-readable artifacts on this domain

    • /llms.txtCurated model-facing index of this site, served at the root path.
    • /llms-full.txtExpanded plain-text corpus of the site's definitions and frameworks.
    • /sitemap.xmlEvery indexable route with image metadata, generated at build time and checked against the router.
    • /feeds/all.xmlDated, machine-readable publication record across guides, dives, and articles.

    All 7 market recordsUpdated 2026-08-21