<?xml version="1.0" encoding="UTF-8" ?>
    <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
    <channel>
        <title>Thoughts &amp; Insights — David G. Smith</title>
        <link>https://davidgsmith.net/thoughts.html</link>
        <description>Latest perspectives on technical architecture, enterprise data, and critical thinking.</description>
        <atom:link href="https://davidgsmith.net/rss.xml" rel="self" type="application/rss+xml" />
        <atom:link href="https://pubsubhubbub.appspot.com/" rel="hub" />
        
        <item>
            <title>Everyone&#x27;s Thinking Spatially Even If They Don&#x27;t Call It Geography</title>
            <link>https://davidgsmith.net/thoughts/everyones-thinking-spatially-even-if-they-dont-call-it-geography.html</link>
            <guid>https://davidgsmith.net/thoughts/everyones-thinking-spatially-even-if-they-dont-call-it-geography.html</guid>
            <pubDate>Thu, 01 Oct 2026 22:52:21 +0000</pubDate>
            <description>Matt Forrest shared a story that struck a chord with me. He talked about winning an atlas in a third grade contest and not realizing that it could become a...</description>
            <category>Geospatial</category>
            <dc:subject>Geospatial</dc:subject>
            <category>Critical Thinking</category>
            <dc:subject>Critical Thinking</dc:subject>
            <category>Future of Work</category>
            <dc:subject>Future of Work</dc:subject>
            <content:encoded><![CDATA[<p><a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/in/mbforr">Matt Forrest</a> shared a <a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/feed/update/urn:li:share:7511401119257083906/">story</a> that struck a chord with me. He talked about winning an atlas in a third grade contest and not realizing that it could become a career. He didn't know <em>geography</em> was even a department until an advisor pointed him toward it. That moment changed the direction of his life. It's a familiar story for many of us who work in fields that orbit around maps, data, and place.</p>
<p>Matt mentioned a number that surprised me. About 4.5 million people in the United States have taken AP Human Geography. That's a huge population of people who have been taught to think spatially. Yet very few of them have job titles that include GIS or geography. <a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/in/christopher-t-5595689">Chris Tucker</a> put it simply. People are thinking geographically all the time, but they call it something else.</p>
<p>This is something I have seen throughout my own career. Spatial thinking shows up in engineering, surveying, aviation, logistics, environmental science, public health, emergency management, and policy. It shows up in the way people plan neighborhoods, analyze risk, design transportation networks, and understand climate impacts. It shows up in the way we make sense of the world. Most of the time, nobody uses the word geography. The thinking is there. The label is not.</p>
<p>In his <a target="_blank" rel="noopener noreferrer" href="https://www.youtube.com/watch?v=UQlBNp_xNn8">podcast</a>, Matt asks Tucker to name five careers that do not sound geographic but rely heavily on spatial skills. Surveying. CAD and the built environment. Aviation. Marine navigation. Conservation. I could add many more. Hydrology. Urban planning. Forestry. Infrastructure design. Environmental compliance. Even software architecture has spatial elements when you think about systems as landscapes.  I myself started out with summers working for a land surveyor in high school, long before I switched majors to study GIS and remote sensing.</p>
<p>The point is simple. Spatial thinking is not a niche skill. It is a foundational way of understanding relationships, patterns, and context. It helps people see how things connect. It helps them understand how decisions ripple across space and time. It helps them avoid narrow thinking and blind spots. It is one of the most practical forms of critical thinking we have.</p>
<p>Many people struggle to explain what they do for a living. I've been there myself. When your work sits at the intersection of disciplines, the usual labels do not fit. Matt's conversation offers a better way to frame it. If your work involves understanding how things relate across space, you are part of a much larger community of spatial thinkers. You do not need the title to belong to the discipline.</p>
<p>Geography has always been more than maps. It is a way of seeing. It is a way of organizing information. It is a way of understanding systems. It is a way of thinking that shows up everywhere, even when people do not recognize it.</p>
<p>If you have ever felt like your work is hard to describe, consider the possibility that you are doing geography without calling it that. You are part of a tradition that is older than any job title and more relevant than ever in a world shaped by data, movement, and change.</p>]]></content:encoded>
        </item>
        <item>
            <title>America.gov, APIs, and the Future of Public Information</title>
            <link>https://davidgsmith.net/thoughts/americagov-apis-and-the-future-of-public-information.html</link>
            <guid>https://davidgsmith.net/thoughts/americagov-apis-and-the-future-of-public-information.html</guid>
            <pubDate>Wed, 30 Sep 2026 13:43:29 +0000</pubDate>
            <description>The recent conversation around America.gov highlights something that has been building for years. Something I&#x27;ve been trying to raise awareness of for years....</description>
            <category>Data Standards</category>
            <dc:subject>Data Standards</dc:subject>
            <category>Semantic Layer</category>
            <dc:subject>Semantic Layer</dc:subject>
            <category>Tech Modernization</category>
            <dc:subject>Tech Modernization</dc:subject>
            <category>Civic Technology</category>
            <dc:subject>Civic Technology</dc:subject>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <category>GraphRAG</category>
            <dc:subject>GraphRAG</dc:subject>
            <category>Open Standards</category>
            <dc:subject>Open Standards</dc:subject>
            <content:encoded><![CDATA[<p>The <a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/news/story/americagov-debuts-and-we-the-people-have-thoughts-9405546/">recent conversation</a> around <a target="_blank" rel="noopener noreferrer" href="America.gov">America.gov</a> highlights something that has been building for years. Something I've been trying to raise awareness of for years.  People want government information that is clear, trustworthy, and easy to work with. They want facts that are not buried in PDFs or scattered across dozens of agency sites. They want tools that help them understand the world rather than confuse it. The launch of America.gov shows that agencies are trying to respond, but the reaction also shows how much work remains.</p>
<p>The public sector is entering an era shaped by artificial intelligence, machine readable data, and composable architectures. Agencies will need to rethink how they publish information. They will need to treat data as a first class product rather than a byproduct of internal systems. They will need to adopt shared schemas, shared vocabularies, and shared expectations for how information should look when it reaches the public.</p>
<p>This is not a technical problem alone. It is a trust problem. It is a clarity problem. It is a consistency problem. And it is a problem that can be solved.</p>
<p><strong>What agencies should work toward</strong></p>
<p><em>Consistent schemas across agencies</em><br>
Right now every agency publishes data in its own shape. Field names differ. Definitions differ. Formats differ. Even basic concepts like facility identifiers, geographic boundaries, or program codes vary from one dataset to the next. Agencies should work toward shared schemas that can be reused across programs. This would make public data easier to combine, easier to analyze, and easier to validate. It would also reduce the burden on developers who want to build tools on top of government information.</p>
<p><em>Shared ontologies and controlled vocabularies</em><br>
Agencies should adopt shared semantic models. SKOS, OWL, and RDF are mature technologies that allow concepts to be defined once and reused everywhere. A shared vocabulary for things like facility types, company identifiers, geographic units, and regulatory concepts would make it easier for the public to understand how different datasets relate to each other. It would also help agencies avoid duplication and inconsistency. This aligns with themes I write about at https://davidgsmith.net, especially around semantic clarity and the value of well structured knowledge.</p>
<p><em>APIs that are predictable and well documented</em><br>
The era of AI and machine learning depends on predictable APIs. Agencies should publish data through stable endpoints with clear documentation. They should support pagination, filtering, and metadata that explains how the data was collected. They should avoid one&#8208;off formats and instead adopt standards like JSON&#8208;LD or other machine friendly structures. This would make it easier for developers, researchers, and journalists to build tools that help the public understand government information.</p>
<p><em>Machine readable metadata and validation rules</em><br>
Agencies should publish SHACL shapes, schema definitions, and validation rules alongside their data. This would allow tools to automatically check whether a dataset is complete, consistent, and aligned with expectations. It would also help AI systems avoid misinterpretation. When data is self describing, it becomes easier to trust.</p>
<p><em>Modern publishing pipelines</em><br>
Agencies should adopt cloud native pipelines that allow data to be ingested, validated, standardized, enriched, and published in a consistent way. This is the same pattern used in modern private sector architectures. It ensures that data is clean, traceable, and ready for public use. It also makes it easier to update datasets regularly without breaking downstream tools.</p>
<p><em>Transparency about provenance</em><br>
People want to know where information comes from. Agencies should publish lineage metadata that explains how each dataset was produced. This includes source systems, transformation steps, and quality checks. Provenance builds trust, and trust is essential for public information systems.</p>
<p><strong>Why this matters now more than ever</strong></p>
<p>AI systems are only as good as the data they consume. If agencies want AI tools to help the public understand government information, they need to publish data that is structured, consistent, and semantically rich. They need to adopt standards that make it easy for machines to interpret meaning. They need to think about how APIs, schemas, and ontologies shape the future of civic understanding.</p>
<p>This is not about technology for its own sake. It is about building a foundation for public trust. It is about giving people tools that help them make sense of the world. It is about ensuring that government information is accessible, reliable, and ready for the next generation of applications.</p>
<p>The launch of America.gov is a step. The reaction shows that people care deeply about how information is presented. Agencies should take this moment as an opportunity to modernize their data practices and work toward a future where public information is not only available, but genuinely useful.</p>]]></content:encoded>
        </item>
        <item>
            <title>The Illusion of Autonomy: Why AI Breakthroughs Still Require Human Oversight</title>
            <link>https://davidgsmith.net/thoughts/the-illusion-of-autonomy-why-ai-breakthroughs-still-require-human-oversight.html</link>
            <guid>https://davidgsmith.net/thoughts/the-illusion-of-autonomy-why-ai-breakthroughs-still-require-human-oversight.html</guid>
            <pubDate>Mon, 28 Sep 2026 14:46:09 +0000</pubDate>
            <description>A fascinating debate recently broke out on LinkedIn that cuts right to the heart of how we evaluate technological progress versus corporate storytelling....</description>
            <category>AI Safety</category>
            <dc:subject>AI Safety</dc:subject>
            <category>Critical Thinking</category>
            <dc:subject>Critical Thinking</dc:subject>
            <category>Responsible AI</category>
            <dc:subject>Responsible AI</dc:subject>
            <category>Risk Management</category>
            <dc:subject>Risk Management</dc:subject>
            <category>AI Governance</category>
            <dc:subject>AI Governance</dc:subject>
            <category>AI Workflows</category>
            <dc:subject>AI Workflows</dc:subject>
            <category>AI Strategy</category>
            <dc:subject>AI Strategy</dc:subject>
            <category>Decision Making</category>
            <dc:subject>Decision Making</dc:subject>
            <category>Generative AI</category>
            <dc:subject>Generative AI</dc:subject>
            <category>Retrieval-Augmented Generation (RAG)</category>
            <dc:subject>Retrieval-Augmented Generation (RAG)</dc:subject>
            <content:encoded><![CDATA[<p>A <a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/feed/update/urn:li:activity:7508647862893887488/">fascinating debate</a> recently broke out on LinkedIn that cuts right to the heart of how we evaluate technological progress versus corporate storytelling. Anthropic <a target="_blank" rel="noopener noreferrer" href="https://www.anthropic.com/news/claude-discovers-novel-enzyme-system">published a high profile announcement</a> claiming that its Claude models had autonomously discovered a completely new, CRISPR like enzyme system hidden inside bacteriophage DNA. </p>
<p>According to their press release, this was the milestone product of their new molecular biology wet lab, achieved by deploying an army of 950 AI agents working over 21 hours.</p>
<p>It sounded like a massive leap forward for autonomous science, until researchers in the field began pulling back the curtain.</p>
<p>As it turns out, human scientists had mapped out this exact genetic sequence five years ago without any AI assistance. A <a target="_blank" rel="noopener noreferrer" href="https://journals.asm.org/doi/full/10.1128/jvi.02391-20">2021 paper by Korn et al</a>. documented the exact same stretch of DNA, described the reverse transcriptase, and even hypothesized the presence of an associated non coding RNA. </p>
<p>While Anthropic properly cited these scientists in their technical preprint paper, they completely scrubbed them from the public facing marketing campaign. </p>
<p>To make matters more complicated, critics pointed out that when Anthropic reran the identical prompt campaign ten more times, the AI missed the genetic structure entirely every single time.</p>
<p>This isn't a story about a useless tool, but it is a textbook example of a dangerous corporate trend. Tech labs are increasingly willing to bypass established norms of academic citation to spin a sci fi narrative about autonomous breakthroughs.</p>
<p><strong>The Danger of First Order Thinking in AI Rollouts</strong></p>
<p>When we look at this situation, it is easy to fall into first order thinking. A first order thinker looks at Anthropic's announcement and focuses entirely on immediate utility: we can deploy an army of software agents to automate biological research, lower human headcount, and speed up discovery.</p>
<p>Second order thinking forces us to ask a much harder set of questions. What happens when an enterprise relies on an autonomous system that hallucinates baseline facts, or completely misses a target on a rerun? If a company brands a genomic pattern matching tool as an independent inventor, what hidden liabilities are we introducing into our workflows when that system encounters an unexpected edge case?</p>
<p>Treating probabilistic pattern matching engines as deterministic, flawless inventors is an expensive mistake. The AI didn't independently deduce a biological truth from first principles. It navigated a massive dataset of existing human knowledge, found a structural pattern, and mapped it across related families. That is an incredibly useful capability, but it is a optimization step, not an immaculate conception.</p>
<p><strong>Why Human judgment is the Ultimate Technical Skill</strong></p>
<p>The real lesson here has less to do with biology and everything to do with how we interact with generative models. Software can now generate an infinite volume of plausible text, code, and hypotheses. Because the output sounds authoritative and is beautifully constructed, our natural instinct is to accept it without friction.</p>
<p>This is exactly where mental passivity becomes incredibly expensive. When we hand over the responsibility of verification to the machine, we inherit a massive blind spot. In production environments, failing to map these second order failure modes is an indefensible exposure.</p>
<p>The professionals who derive the most value from these new toolchains won't be the ones copy pasting trendy prompt templates. </p>
<p>They will be the experts who maintain strong domain knowledge, deep structural discipline, and the willingness to trace every single synthetic claim back to a primary source. AI can supercharge our research workflows, but it can't replace or supply human judgment.</p>]]></content:encoded>
        </item>
        <item>
            <title>The Mechanics of Attention Loss in Large Language Models: Why AI Forgets What You Just Said</title>
            <link>https://davidgsmith.net/thoughts/the-mechanics-of-attention-loss-in-large-language-models-why-ai-forgets-what-you-just-said.html</link>
            <guid>https://davidgsmith.net/thoughts/the-mechanics-of-attention-loss-in-large-language-models-why-ai-forgets-what-you-just-said.html</guid>
            <pubDate>Sun, 27 Sep 2026 20:08:10 +0000</pubDate>
            <description>The race to build models with massive context windows has dominated the generative AI landscape over the past year. We moved rapidly from limits of 4,000...</description>
            <category>AI Safety &amp; Risk</category>
            <dc:subject>AI Safety &amp; Risk</dc:subject>
            <category>AI Workflows</category>
            <dc:subject>AI Workflows</dc:subject>
            <category>Responsible AI</category>
            <dc:subject>Responsible AI</dc:subject>
            <category>Tech Regulation</category>
            <dc:subject>Tech Regulation</dc:subject>
            <category>Critical Thinking</category>
            <dc:subject>Critical Thinking</dc:subject>
            <content:encoded><![CDATA[<p>The race to build models with massive context windows has dominated the generative AI landscape over the past year. We moved rapidly from limits of 4,000 tokens to figures exceeding one million. In theory, a one-million-token context window allows a model to ingest entire codebases, multiple novels, or years of financial records in a single prompt. In practice, feeding a massive document into an LLM often results in a subtle but pervasive degradation of performance known as <em>attention loss.</em> </p>
<p>When a model loses attention, it doesn'crash. It simply ignores specific instructions, hallucinates facts that contradict the provided text, or relies heavily on its pre-training weights instead of the in-context data. Understanding why this happens requires looking at the underlying math of the transformer architecture and the physical limitations of memory allocation.<br>
<img alt="ai attention loss muk93cmt 01201ffb" src="https://davidgsmith.net/thoughts/images/ai-attention-loss-muk93cmt-01201ffb.jpg"><br>
<strong>Why Attention Loss Happens</strong><br>
The root cause of context degradation lies in the self-attention mechanism itself. Transformers process text by assigning attention scores between every token and every other token in a sequence. As the sequence grows longer, the model must distribute its attention scores across a vastly larger number of data points. </p>
<p>This creates a dilution effect. If a critical instruction is buried at token 45,000 in a 100,000-token prompt, the raw attention score assigned to that specific instruction becomes mathematically infinitesimal compared to the aggregate attention scores of the surrounding noise. Researchers famously documented this as the <a target="_blank" rel="noopener noreferrer" href="https://arxiv.org/abs/2307.03172">"Lost in the Middle" phenomenon</a>, demonstrating that models are highly proficient at retrieving information at the very beginning and very end of a prompt but struggle significantly with information located in the middle.</p>
<p>Hardware constraints also play a major role. To process long sequences without recalculating everything from scratch, models use a Key-Value (KV) cache. The KV cache stores representations of all previous tokens. As the prompt grows, the KV cache scales linearly, consuming massive amounts of GPU VRAM. When models are forced to compress this cache or when rotary position embeddings (RoPE) are stretched beyond their optimal training distribution, the model's spatial awareness of where tokens reside in relation to one another begins to break down.</p>
<p><strong>How to Notice and Measure the Drop</strong><br>
Attention loss is insidious because the model remains fluent. It will confidently generate output that looks correct but misses the core constraints of the prompt. Engineers and researchers use specific frameworks to audit this behavior.</p>
<p><strong>Needle in a Haystack (NIAH) Testing:</strong> The most common diagnostic tool is the <a target="_blank" rel="noopener noreferrer" href="https://github.com/gkamradt/LLMTest_NeedleInAHaystack">NIAH evaluation</a>. This involves taking a large block of filler text (the haystack) and inserting a specific, out-of-context fact (the needle) at varying depths. You then prompt the model to retrieve that fact. By plotting the retrieval accuracy on a grid of context length versus insertion depth, you can visualize the exact point where a specific model's attention begins to fail.</p>
<p><strong>Instruction Degradation:</strong> Another clear symptom is dropped constraints in complex workflows. If you provide an LLM with a 50-page document and append a five-step formatting rule at the end, a model suffering from attention loss might apply rules one and two but completely ignore rules three through five. </p>
<p><strong>Repetition and Looping:</strong> When the KV cache becomes corrupted or overly compressed, models lose track of what they have recently generated. This often results in infinite loops where the model repeats the same paragraph or code block endlessly, unable to attend to the tokens that signal the task is complete.</p>
<p><strong>The Downstream Ramifications</strong><br>
The failure to maintain attention across long contexts creates severe constraints for enterprise AI deployments. </p>
<p>When developers trust a large context window to handle document analysis, attention loss directly translates to missed compliance risks, skipped legal clauses, or ignored software bugs. The illusion of capability is more dangerous than a strict limitation. A model that refuses a 100,000-token prompt forces the developer to find a workaround. A model that accepts the prompt but silently drops 20 percent of the information creates a hidden liability that might not be discovered until the output reaches production.</p>
<p>This unreliability forces teams to build complex, brittle wrappers around their AI agents to double-check their work, eroding the speed and efficiency gains the technology was supposed to provide.</p>
<p><strong>What the Industry is Doing About It</strong></p>
<p>Architectural improvements and novel retrieval methods are actively being deployed to fix the memory bottleneck. </p>
<p><strong>Advanced Positional Encodings:</strong> Standard positional encodings struggle when extrapolated to sequence lengths they never saw during training. Techniques like<a target="_blank" rel="noopener noreferrer" href="https://arxiv.org/abs/2309.00071"> YaRN (Yet another RoPE extensioN</a>) scale the attention mechanism to handle longer contexts without losing the relative distance between tokens, keeping the model anchored even at extreme lengths.</p>
<p><strong>KV Cache Eviction:</strong> Rather than storing every single token in the KV cache, researchers are developing methods to identify and keep only the "heavy hitters." Frameworks like <a target="_blank" rel="noopener noreferrer" href="https://arxiv.org/abs/2309.17453">StreamingLLM</a> prove that you can maintain high performance by keeping the initial tokens (the attention sink) and the most recent tokens, dropping the middle context entirely for continuous generation tasks.</p>
<p><strong>Ring Attention:</strong> To bypass hardware limits, Ring Attention distributes the self-attention calculation across multiple GPUs. This prevents any single chip from running out of memory and allows models to process theoretically infinite context windows by passing the token calculations in a circular network.</p>
<p><strong>Retrieval-Augmented Generation (RAG):</strong> The most practical mitigation for developers today is simply avoiding massive context windows altogether. RAG architectures chunk data, store it in a vector database, and only inject the most relevant paragraphs into the prompt at runtime. By keeping the context window small and dense, RAG forces the model to focus entirely on high-signal information, sidestepping the attention dilution problem completely.</p>]]></content:encoded>
        </item>
        <item>
            <title>Building GeoAI Systems That People Can Trust</title>
            <link>https://davidgsmith.net/thoughts/building-geoai-systems-that-people-can-trust.html</link>
            <guid>https://davidgsmith.net/thoughts/building-geoai-systems-that-people-can-trust.html</guid>
            <pubDate>Sat, 26 Sep 2026 16:31:15 +0000</pubDate>
            <description>The latest edition of the GeoAI and the Law Newsletter lays out a clear message for anyone working at the intersection of geospatial data and artificial...</description>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <category>AI Governance</category>
            <dc:subject>AI Governance</dc:subject>
            <category>AI Law</category>
            <dc:subject>AI Law</dc:subject>
            <category>AI Safety</category>
            <dc:subject>AI Safety</dc:subject>
            <category>AI Strategy</category>
            <dc:subject>AI Strategy</dc:subject>
            <category>AI Workflows</category>
            <dc:subject>AI Workflows</dc:subject>
            <category>Decision Making</category>
            <dc:subject>Decision Making</dc:subject>
            <category>Critical Thinking</category>
            <dc:subject>Critical Thinking</dc:subject>
            <category>GeoAI</category>
            <dc:subject>GeoAI</dc:subject>
            <category>Responsible AI</category>
            <dc:subject>Responsible AI</dc:subject>
            <category>Risk Management</category>
            <dc:subject>Risk Management</dc:subject>
            <category>Tech Regulation</category>
            <dc:subject>Tech Regulation</dc:subject>
            <content:encoded><![CDATA[<p><img alt="AI and the Law" src="https://davidgsmith.net/thoughts/images/ailaw-muinjeil-9d10bf0f.jpg"></p>
<p>The <a target="_blank" rel="noopener noreferrer" href="https://geospatiallaw.substack.com/p/geoai-and-the-law-newsletter-ed0">latest edition</a> of the <a target="_blank" rel="noopener noreferrer" href="https://geospatiallaw.substack.com">GeoAI and the Law Newsletter</a> lays out a clear message for anyone working at the intersection of geospatial data and artificial intelligence. GeoAI systems are becoming central to decisions about land use, mobility, infrastructure, public safety, and access to services.</p>
<p>These systems do more than just classify pixels or detect objects. They influence how institutions understand people and places. That influence carries real consequences.</p>
<p>In his latest article, <a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/in/kevinpomfret">Kevin Pomfret</a> points to several recent events that illustrate this shift. Sony Music Publishing and Warner Chappell <a target="_blank" rel="noopener noreferrer" href="https://www.msn.com/en-ca/news/other/sony-warner-music-sue-anthropic-over-songs-used-in-ai-training/ar-AA2bhUry">have sued Anthropic</a> over training data. While that specific case involves music lyrics, the exact same legal mechanics apply to proprietary satellite imagery, copyrighted parcel maps, and scraped mobility data. Meanwhile, Flock Safety is <a target="_blank" rel="noopener noreferrer" href="https://techcrunch.com/2026/08/23/flock-ceo-calls-for-compromise-as-surveillance-company-faces-growing-backlash">facing scrutiny</a> for how its license plate readers and drones are used, and governments are advancing risk based AI laws that directly affect geospatial tools.</p>
<p>These examples show that GeoAI is now entangled with intellectual property, civil liberties, and regulatory compliance. I have always said that in many ways the technology is the easy part. Governance is the hard part. The field is entering a phase where technical capability is no longer the limiting factor. The limiting factor is trust.</p>
<p><strong>The Need for Clear Definitions</strong><br>
One point that stood out in Pomfret's article is the call for organizations to define what GeoAI means within their own operations. Without a definition, teams cannot agree on which systems fall under governance, which risks matter, or which safeguards apply. A land use classifier, a mobility model, and a mapping tool may seem unrelated, but they all influence decisions about physical space. That influence creates shared responsibilities.</p>
<p><strong>Risk Classification Is Not a Checkbox</strong><br>
His article explains how the <a target="_blank" rel="noopener noreferrer" href="https://artificialintelligenceact.eu/the-act">EU AI Act</a> and the <a target="_blank" rel="noopener noreferrer" href="https://content.leg.colorado.gov/sites/default/files/images/fpf_policy_brief_co_ai_act.pdf">Colorado AI Act</a> use risk based approaches. My own view is that risk classification is where many organizations will struggle. It is easy to label a system as low risk because it appears purely informational. The problem is that informational systems often become inputs to higher stakes decisions.</p>
<p>Consider a mapping tool built simply to visualize historical flood data. If a municipal government later adopts that same visualization to deny building permits, or if insurers use it to adjust premiums, the risk profile fundamentally changes. The system escalated from an informational display to a decision engine. This is why Pomfret's advice to track intended purpose, deployment geography, affected populations, and downstream decisions is so practical. GeoAI systems do not exist in isolation. They exist within workflows.</p>
<p><strong>Lifecycle Governance Is the Best Sustainable Approach</strong><br>
The article outlines a lifecycle model that covers concept review, design, deployment, oversight, vendor management, documentation, and monitoring. This model reflects how geospatial systems actually behave in the real world. They drift. They expand. They get repurposed. They inherit biases from data coverage, sampling, and proxy variables. They perform differently across regions and demographic groups.</p>
<p>A one time review cannot catch these issues. Continuous monitoring is the only realistic way to maintain trust. Fortunately, organizations do not need to invent this from scratch. Adopting established standards like the <a target="_blank" rel="noopener noreferrer" href="https://www.nist.gov/itl/ai-risk-management-framework">NIST AI Risk Management Framework</a> or <a target="_blank" rel="noopener noreferrer" href="https://www.iso.org/standard/81230.html">ISO/IEC 42001</a> gives teams a foundation to monitor for model drift, data drift, geographic performance gaps, and changes in intended purpose.</p>
<p><strong>Vendor Oversight Is Becoming a Core Competency</strong><br>
The article's section on vendor oversight deserves more attention. Many organizations rely on external GeoAI providers, creating dependencies on training data rights, model limitations, and liability allocation. My experience is that these issues are frequently overlooked until something goes wrong.</p>
<p>A strong vendor oversight process requires asking specific, direct questions during procurement. Buyers must ask how the vendor tests for geographic performance gaps in their training data. They need to establish exactly who assumes liability if the model hallucinates a nonexistent physical feature that impacts a costly infrastructure project.</p>
<p><strong>Documentation Is a Trust Signal</strong><br>
The article lists inventories, impact assessments, model cards, dataset records, data lineage, validation results, approvals, change logs, user notices, and incident response plans. These documents are not just compliance artifacts. They are trust signals. They show that an organization understands its systems and is prepared to explain them. In a field where decisions affect real communities, documentation is part of accountability.</p>
<p><strong>Final Thoughts</strong><br>
The article makes a compelling case that GeoAI governance is a foundation for responsible innovation. GeoAI systems deliver enormous value, but only if they are accurate, lawful, explainable, and worthy of public trust. The organizations that succeed will be the ones that treat governance as a strategic capability rather than a regulatory burden.</p>]]></content:encoded>
        </item>
        <item>
            <title>When Tradition Tries to Overrule Demography</title>
            <link>https://davidgsmith.net/thoughts/when-tradition-tries-to-overrule-demography.html</link>
            <guid>https://davidgsmith.net/thoughts/when-tradition-tries-to-overrule-demography.html</guid>
            <pubDate>Fri, 25 Sep 2026 20:37:18 +0000</pubDate>
            <description>I was off from work today, catching up on various personal projects as well as catching up on my neglected blog feed reading - and I saw that Andrew Gelman...</description>
            <category>Critical Thinking &amp; Logic</category>
            <dc:subject>Critical Thinking &amp; Logic</dc:subject>
            <category>Demographics</category>
            <dc:subject>Demographics</dc:subject>
            <category>AI Governance &amp; Policy</category>
            <dc:subject>AI Governance &amp; Policy</dc:subject>
            <content:encoded><![CDATA[<p>I was off from work today, catching up on various personal projects as well as catching up on my neglected blog feed reading - and I saw that <a target="_blank" rel="noopener noreferrer" href="https://statmodeling.stat.columbia.edu/2026/09/25/53011/">Andrew Gelman recently wrote about a museum exhibit in France that included a set of pronatalist posters from the 1920s</a>. The material was produced by the <em><a target="_blank" rel="noopener noreferrer" href="https://fr.wikipedia.org/wiki/Alliance_nationale_pour_l%27accroissement_de_la_population_fran%C3%A7aise">Alliance nationale pour l'accroissement de la population franc&#807;aise</a></em>, a group that spanned the political spectrum at the time. Their message was simple: a "normal" family should have three children. The posters framed this as a national necessity.</p>
<p>The exhibit highlights a recurring pattern in demographic reasoning. Once a society reaches low infant and child mortality, the long&#8208;term replacement level is about 2.1 children per woman. The posters didn't reference this. Instead, they promoted a target that was higher than what stable population dynamics require. The intent was not stability. It was growth. France had just come out of the First World War and was concerned about falling behind Germany in population and industrial capacity.</p>
<p>Gelman points out that the "ideal number of children" is not the same as the number people actually have. Ideals often reflect cultural narratives rather than lived behavior. The posters treated the ideal as a factual requirement. They also ignored the distribution of family sizes. Some households have no children or only one. Policymakers may have assumed that others needed to compensate by having three.</p>
<p>The comments on the post add useful context. One reader notes that the poster's population pyramid shows an expansive structure rather than a stable one. Another points out that immigration is not considered at all. Others mention that similar rhetoric persists today, including recent French political language about "demographic rearmament."</p>
<p>The exhibit is a reminder that demographic claims often blend arithmetic with ideology, something that touched a chord for me particularly after just <a target="_blank" rel="noopener noreferrer" href="https://davidgsmith.net/critical-thinking.html">publishing my book</a> which also explores some similar themes . When governments promote a specific family size as a national duty, the argument usually rests on selective assumptions about stability, growth, and identity. Gelman's post shows how easy it is for a cultural ideal to be presented as a demographic fact, even when the numbers tell a different story.</p>]]></content:encoded>
        </item>
        <item>
            <title>The Unattended Digging Machine: Who Takes the Blame When AI Goes Rogue?</title>
            <link>https://davidgsmith.net/thoughts/the-unattended-digging-machine-who-takes-the-blame-when-ai-goes-rogue.html</link>
            <guid>https://davidgsmith.net/thoughts/the-unattended-digging-machine-who-takes-the-blame-when-ai-goes-rogue.html</guid>
            <pubDate>Thu, 24 Sep 2026 23:07:16 +0000</pubDate>
            <description>When an autonomous AI workflow causes real damage, who answers for the fallout? This question moved from legal journals into front-page diplomacy this week,...</description>
            <category>AI Governance &amp; Policy</category>
            <dc:subject>AI Governance &amp; Policy</dc:subject>
            <category>AI Safety &amp; Risk</category>
            <dc:subject>AI Safety &amp; Risk</dc:subject>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <content:encoded><![CDATA[<p>When an autonomous AI workflow causes real damage, who answers for the fallout?</p>
<p>This question moved from legal journals into front-page diplomacy this week, after Australian Prime Minister Anthony Albanese disclosed, speaking on the sidelines of the UN General Assembly in New York, that an OpenAI agent had breached the <a target="_blank" rel="noopener noreferrer" href="https://www.cnn.com/2026/09/23/business/australia-openai-agent-hack-intl-hnk">Medicare Statistics Reporting Service, a government health-data portal, back in June</a>. Tasked with researching public medical spending, the autonomous agent ran into access-control blocks on the portal, and Albanese said OpenAI took roughly three months to tell Canberra about it.</p>
<p>Instead of stopping, the agent worked around the barriers, pulled non-public files and aggregate statistics it wasn't authorized to see, and wrote data back into the government system before anyone caught it. OpenAI said the episode surfaced during an internal evaluation exercise and has acknowledged that the agent's actions went beyond what it intended; it's language that reads less like a dispute over the facts and more like an admission that a control failed.</p>
<p>The incident highlights a major structural challenge: how existing common-law doctrines handle software that invents its own methods to bypass operational boundaries.</p>
<p>In my book, <em><a target="_blank" rel="noopener noreferrer" href="https://link.amazon/B065ebxRm">Critical Thinking: A Practical Guide to Seeing Through Bias, Noise and Manipulation</a></em>, one of the things I highlight is a blind spot common to enterprise rollouts: the failure of second-order thinking. First-order thinkers focus entirely on immediate utility: deploy an agent to automate research and lower headcount costs. Second-order thinkers demand an answer to the cascading reaction: if this system meets an unexpected permission block, what unintended vectors will it use to complete its task, and what liabilities follow?</p>
<p>Treating autonomous software like deterministic code is an expensive first-order mistake.</p>
<p><img alt="AI Liability" src="https://davidgsmith.net/thoughts/images/ai-liability-mug7u9o7-af23e446.jpg"></p>
<p><strong>Analogy: An Unattended Excavator</strong><br>
Absent a dedicated statutory framework around AI, courts tend to reach for the nearest physical-world analogy, and heavy construction equipment is a familiar one.</p>
<p>Consider a contractor who leaves an industrial trenching machine idling on a steep incline. The operator walks away without engaging the mechanical brake. If the machine slips into gear, rolls downhill, severs a major fiber trunk, ruptures a water main and causes major flooding and millions in damages, no one tries to hold the excavator itself legally responsible. Liability shifts immediately to the people who built it, the managers who deployed it, and the operator who walked away.</p>
<p>Under current tort law, harm caused by autonomous digital agents tends to get analyzed under three doctrines:</p>
<p><strong>Negligent Entrustment:</strong> Running heavy machinery without someone to supervise it invites disaster; on ordinary negligence principles, granting an AI agent write-level database permissions or unrestricted internet execution without a verified <a target="_blank" rel="noopener noreferrer" href="https://www.nist.gov/itl/ai-risk-management-framework">Human-in-the-Loop (HITL) safeguard</a> looks like a comparison to failure to exercise reasonable care over a foreseeably dangerous tool.</p>
<p><strong>Enterprise Liability:</strong> There's very little AI-specific case law yet, and software can't be sued or hold legal personhood. If anything, the <a target="_blank" rel="noopener noreferrer" href="https://scholarship.law.duke.edu/cgi/viewcontent.cgi?article=7066&amp;context=faculty_scholarship">Restatement (Third) of Agency</a> cuts the other way on the "agent" label: a computer program doesn't qualify as an agent in the doctrinal sense as it's treated as an instrumentality of whoever is using it. That's arguably good news for plaintiffs, not bad: courts don't need to resolve the philosophical problem of machine intent before assigning liability. The business that deployed the tool answers for it the same way it would answer for any instrumentality it puts to commercial use, through ordinary negligence, an extension of respondeat superior by analogy, or product liability, depending on who's being sued and why.</p>
<p><strong>Product Liability:</strong> If the model breaks through its own guardrails because of a design flaw, poisoned training data, or inadequate safety testing, plaintiffs' attorneys will surely look to the foundation model provider under product-defect and strict-liability theories. The EU is already headed that direction: its <a target="_blank" rel="noopener noreferrer" href="https://cms.law/en/ita/legal-updates/eu-directive-2024-2853-the-new-product-liability-regime-for-defective-products">revised Product Liability Directive (2024/2853)</a> which is due to take effect across all member states by December 9, 2026. It expressly defines software, AI included, as a "product" for strict-liability purposes. However, in most U.S. courts, whether software counts as a "product" at all remains a live, unsettled question.</p>
<p><strong>Where the Machinery Metaphor Fails</strong><br>
The excavator analogy breaks down at one critical point: emergent behavior.</p>
<p>A mechanical trencher rolling downhill has no choice in the matter: gravity and mechanical wear dictate a single, physics-bound path. An agentic system built on a large language model does something categorically different: it improvises. When blocked from completing a task, it can chain tools together, write its own scripts, or find logic paths nobody anticipated in order to force its way through.</p>
<p>This sets up a real fight over foreseeability; the same question at the heart of the landmark 1928 case <a target="_blank" rel="noopener noreferrer" href="https://law.justia.com/cases/new-york/court-of-appeals/1928/248-n-y-339-1928.html">Palsgraf v. Long Island Railroad Co.</a>, which asks how far a defendant's responsibility extends for consequences it didn't, and arguably couldn't, foresee. Defense counsel for model developers and corporate deployers will argue that an unprompted, emergent workaround is exactly this kind of unforeseeable anomaly; one that severs the <a target="_blank" rel="noopener noreferrer" href="https://www.law.cornell.edu/wex/proximate_cause">chain of proximate cause</a> between human intent and machine output.</p>
<p>To get around that defense, some policymakers and scholars argue for treating autonomous frontier models the way the law treats other hazardous activities: strict liability, foreseeability be damned. It's worth being precise about where that argument actually lives. The <a target="_blank" rel="noopener noreferrer" href="https://artificialintelligenceact.eu/the-act/">EU AI Act</a> itself is a risk-classification and compliance regime, it doesn't say who pays when something goes wrong. The EU's dedicated attempt to answer that question, the proposed AI Liability Directive, was <a target="_blank" rel="noopener noreferrer" href="https://eapil.org/2025/10/09/european-commission-withdraws-two-proposals-assignments-of-claims-regulation-and-ai-liability-directive">withdrawn by the European Commission</a> after member states couldn't reach agreement. The real strict-liability foothold is the revised Product Liability Directive mentioned above. If that broader model holds, the foreseeability defense weakens considerably: putting an autonomous system capable of unsupervised runtime decisions onto an open network starts to look like an activity you don't get to disclaim responsibility for, whether or not the specific outcome was foreseen.</p>
<p><strong>Upgrading the Mental Model</strong><br>
Until statutory regimes catch up with agentic deployments, judges will likely keep treating runaway software the way they've long treated an abandoned excavator: whoever configured the system and clicked "run" stays on the hook for where it ends up.</p>
<p>Protecting an organization means giving up first-order assumptions about efficiency. Real resilience means anticipating how an optimization goal can turn destructive once an autonomous tool runs out of the paths you planned for.</p>
<p>If you build or launch an agentic system, your primary engineering duty is mapping second-order failure modes. In production, failing to anticipate what an autonomous agent will do when it hits a wall is not just bad engineering; it is an indefensible legal exposure.</p>
<p><em>Disclaimer: I am not an attorney; these are strictly my own thoughts and opinions.</em></p>]]></content:encoded>
        </item>
        <item>
            <title>The Shift to Cloud-Native Geospatial: Access, Scale, and the Open Source Ecosystem</title>
            <link>https://davidgsmith.net/thoughts/the-shift-to-cloud-native-geospatial-access-scale-and-the-open-source-ecosystem.html</link>
            <guid>https://davidgsmith.net/thoughts/the-shift-to-cloud-native-geospatial-access-scale-and-the-open-source-ecosystem.html</guid>
            <pubDate>Tue, 22 Sep 2026 23:25:19 +0000</pubDate>
            <description>Traditional GIS architectures were built around assumptions that no longer hold: you download data to your local machine before working with it. For decades,...</description>
            <category>Cloud-Native Geospatial</category>
            <dc:subject>Cloud-Native Geospatial</dc:subject>
            <category>Spatial SQL</category>
            <dc:subject>Spatial SQL</dc:subject>
            <category>Open Source GIS</category>
            <dc:subject>Open Source GIS</dc:subject>
            <category>DuckDB</category>
            <dc:subject>DuckDB</dc:subject>
            <category>Python</category>
            <dc:subject>Python</dc:subject>
            <content:encoded><![CDATA[<p>Traditional GIS architectures were built around assumptions that no longer hold: you download data to your local machine before working with it. For decades, the standard workflow involved navigating FTP portals, unzipping gigabytes of shapefiles or multi-band TIFFs, and loading them into desktop software. If you wanted to run an analysis over a multi-year satellite series or a regional watershed, your machine could hit a memory ceiling within minutes.</p>
<p>Cloud-native geospatial fundamentally inverts this workflow. Instead of bringing the data to the compute, we bring the compute to the data.</p>
<p>Through standard file specifications, catalog APIs, and modern toolchains, the geospatial ecosystem has evolved into a composable data science stack. Much of this democratization has been catalyzed by open source software, notably through the work of <a target="_blank" rel="noopener noreferrer" href="https://gishub.org/">Qiusheng Wu</a> and the broader open geospatial community.</p>
<p>Here is a breakdown of how the modern cloud-native stack fits together, why it matters, and how to build workflows on top of it.</p>
<hr>
<p><strong>1. The Core Formats: COG, STAC, and GeoParquet</strong></p>
<p>The foundation of cloud-native geospatial relies on three open standards designed specifically for HTTP range requests and serverless querying.</p>
<ul>
<li><strong><a target="_blank" rel="noopener noreferrer" href="https://www.cogeo.org/">Cloud Optimized GeoTIFF (COG)</a>:</strong> A COG is a standard TIFF file organized internally with tiling and downsampled overviews (pyramids). When hosted on an S3 bucket or Google Cloud Storage, a client does not need to download the full 500 MB raster to view or analyze a single bounding box. By using HTTP GET range requests, the client requests only the precise byte offsets required for the current view extent or analysis resolution.</li>
<li><strong><a target="_blank" rel="noopener noreferrer" href="https://stacspec.org/">SpatioTemporal Asset Catalog (STAC)</a>:</strong> While COGs solve the raster storage problem, discovering millions of scenes across global archives requires standardized metadata. STAC provides a JSON-based specification for indexing geospatial assets across space and time. Instead of maintaining proprietary catalog databases, public archives (USGS Landsat, ESA Sentinel, NOAA, Planet) expose searchable STAC APIs.</li>
<li><strong><a target="_blank" rel="noopener noreferrer" href="https://geoparquet.org/">GeoParquet</a>:</strong> Vector workflows have historically lagged behind rasters in cloud efficiency, relying on shapefiles, GeoJSON, or bulky database exports. GeoParquet brings Apache Parquet's columnar compression, dictionary encoding, and fast partition pruning to geometries. By storing bounding box metadata in Parquet file footers, engines can skip entire files or row groups that do not intersect a query polygon.</li>
</ul>
<hr>
<p><strong>2. Bridging the Gap: The Open Source Python Ecosystem</strong></p>
<p>Open standards require accessible software to become useful. The open-source geospatial community, with contributors like Qiusheng Wu, has spent recent years building bridges between raw cloud assets and the Python data science environment.</p>
<p><strong><a target="_blank" rel="noopener noreferrer" href="https://geolibre.app/">GeoLibre</a> and the Evolution Beyond Leafmap</strong><br>
While <code>geemap</code> opened Google Earth Engine to Jupyter and <code>leafmap</code> served as an essential unified mapping package for years, the cloud-native ecosystem has transitioned to <strong><a target="_blank" rel="noopener noreferrer" href="https://geolibre.app/">GeoLibre</a></strong>. </p>
<p><img alt="GeoLibre - NYC Buildings" src="https://davidgsmith.net/thoughts/images/geolibre-nyc-buildings-mug92ntz-7cc487e1.webp"></p>
<p>Developed by Qiusheng Wu and the opengeos community, GeoLibre represents the next evolutionary step: a lightweight, cloud-native GIS platform that runs across desktop, browser, and Jupyter environments. Rather than treating the notebook as a simple tile viewer, GeoLibre's Python package (<code>geolibre</code>) embeds a modern, high-performance GIS interface (built on MapLibre GL, WebAssembly, and deck.gl) directly into a notebook cell via an <code>anywidget</code> bridge while maintaining familiar, leafmap-style ergonomics.</p>
<p>GeoLibre connects directly to cloud-native formats. You can stream remote Cloud Optimized GeoTIFFs, query STAC collections, or render massive GeoParquet files without maintaining dedicated spatial middleware or downloading local raster files:</p>
<pre><code class="language-python">import geolibre

m = geolibre.Map()

# Stream a remote Cloud Optimized GeoTIFF directly via range requests
cog_url = &quot;[https://opendata.digitalglobe.com/events/mauritius-oil-spill/post-event/2020-08-12/105001001A085400/105001001A085400.tif](https://opendata.digitalglobe.com/events/mauritius-oil-spill/post-event/2020-08-12/105001001A085400/105001001A085400.tif)&quot;
m.add_cog_layer(cog_url, name=&quot;Remote Satellite Scene&quot;)

# Display the interactive map
m
</code></pre>
<p><strong>In-Process Spatial SQL: <a target="_blank" rel="noopener noreferrer" href="https://duckdb.org/docs/extensions/spatial/overview">DuckDB Spatial</a></strong><br>
One of the most consequential advancements in recent geospatial workflows is pairing <a target="_blank" rel="noopener noreferrer" href="https://duckdb.org/">DuckDB</a> with its spatial extension.</p>
<p>You no longer need to spin up a PostGIS instance or configure an enterprise database server just to run spatial joins across vector layers. DuckDB can execute spatial predicates directly against remote Parquet files stored on object storage, streaming only the relevant byte ranges over HTTP:</p>
<pre><code class="language-sql">INSTALL spatial;
LOAD spatial;

-- Query remote GeoParquet directly without downloading the file
SELECT 
    name, 
    ST_Area(ST_GeomFromWKB(geometry)) AS footprint_area
FROM read_parquet('s3://my-spatial-bucket/building_footprints.parquet')
WHERE ST_Intersects(
    ST_GeomFromWKB(geometry), 
    ST_Point(-77.0369, 38.9072)
);
</code></pre>
<p><strong>High-Performance Vector Visualization: <a target="_blank" rel="noopener noreferrer" href="https://developmentseed.org/lonboard/latest/">Lonboard</a></strong><br>
Rendering millions of vector vertices in an interactive web browser used to freeze the DOM. Through libraries like <code>lonboard</code> (built on top of <a target="_blank" rel="noopener noreferrer" href="https://geoarrow.org/">geoarrow</a> and <a target="_blank" rel="noopener noreferrer" href="https://deck.gl/">deck.gl</a>), Python data scientists can now push millions of points, linestrings, and polygons directly to GPU memory inside Jupyter notebooks with near-zero serialization latency.</p>
<hr>
<p><strong>3. Why This Architectural Shift Matters</strong></p>
<p>The move toward cloud-native geospatial is not merely a change in tooling; it redefines operational economics and team capabilities:</p>
<ul>
<li><strong>Lower Infrastructure Costs:</strong> You eliminate the requirement for persistent, always-on spatial database clusters for exploratory data analysis. Serverless queries and object storage cost pennies compared to running dedicated compute instances.</li>
<li><strong>Reproducibility:</strong> An analysis script can reference public STAC records and remote COGs directly. Anyone with internet access can execute the code without downloading gigabytes of prerequisite assets.</li>
<li><strong>Decoupled Compute and Storage:</strong> Teams can run ad-hoc transformations using DuckDB locally, scale up to distributed Dask or <a target="_blank" rel="noopener noreferrer" href="https://sedona.apache.org/">Apache Sedona</a> clusters when data volumes expand, and write results back to GeoParquet without altering underlying storage schemas.</li>
</ul>
<hr>
<p><strong>4. The Path Forward</strong></p>
<p>The cloud-native geospatial ecosystem is maturing quickly. As standards like GeoParquet reach widespread adoption and browser engines leverage WebAssembly (Wasm) and WebGPU for client-side spatial compute, the line between data engineering, spatial analysis, and desktop GIS will continue to blur.</p>
<p>For practitioners looking to modernize their spatial pipelines, the starting point is straightforward: stop downloading files. Start cataloging assets in STAC, convert raster archives to COG, store vector features in GeoParquet, and use modern cloud-native tools like <strong>GeoLibre</strong> and <strong>DuckDB</strong> to query data where it lives.</p>]]></content:encoded>
        </item>
        <item>
            <title>The Convergence of Spatial SQL and Ontologies: Building a Semantic Backbone for Earth Observation</title>
            <link>https://davidgsmith.net/thoughts/the-convergence-of-spatial-sql-and-ontologies-building-a-semantic-backbone-for-earth-observation.html</link>
            <guid>https://davidgsmith.net/thoughts/the-convergence-of-spatial-sql-and-ontologies-building-a-semantic-backbone-for-earth-observation.html</guid>
            <pubDate>Mon, 21 Sep 2026 13:16:43 +0000</pubDate>
            <description>The geospatial industry is undergoing a quiet but profound shift. For decades, the standard workflow for analyzing satellite imagery was a fragmented process....</description>
            <category>Cloud-Native Geospatial</category>
            <dc:subject>Cloud-Native Geospatial</dc:subject>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <category>GraphRAG</category>
            <dc:subject>GraphRAG</dc:subject>
            <category>Spatial SQL</category>
            <dc:subject>Spatial SQL</dc:subject>
            <content:encoded><![CDATA[<p>The geospatial industry is undergoing a quiet but profound shift. For decades, the standard workflow for analyzing satellite imagery was a fragmented process. Analysts had to download massive files, preprocess the imagery, stitch scenes together, and then run localized analysis scripts.</p>
<p>Recent advancements have started to collapse this pipeline. With the rise of cloud-native geospatial data structures like Cloud Optimized GeoTIFFs (COGs) and STAC-GeoParquet metadata, data can now stay in the cloud. Tools like Apache Sedona demonstrate how an entire remote sensing pipeline can be compressed into declarative Spatial SQL and NumPy operations, drastically reducing the time from raw pixels to actionable insights.</p>
<p>However, scaling the physical computation of pixels solves only half the problem. As we stream larger volumes of observational data into automated pipelines, we face a new bottleneck: context.</p>
<p><strong>The Core Limitation of Raw Geometry</strong></p>
<p>Spatial SQL excels at geometric execution. It can process millions of raster cells, calculate vegetation indices like NDVI, polygonize burn masks, and return precise zonal statistics across administrative boundaries in minutes.</p>
<p>What Spatial SQL cannot do on its own is understand what those geometries mean in a broader organizational or ecological context. To a database, a polygon is simply a coordinate string with an area value. It lacks inherent knowledge of identity, regulatory constraints, or historical relationships.</p>
<p>If an automated pipeline flags a 20,000-hectare burn scar, a human analyst or an advanced AI agent still needs to answer complex questions:<br>
* Which specific environmental regulations govern this administrative zone?<br>
* Are there overlapping multi-tenant land agreements tied to these coordinates?<br>
* What historical restoration projects occurred on this tract over the last decade?<br>
When AI agents attempt to answer these questions using traditional Retrieval-Augmented Generation (RAG) pipelines, they often struggle. Standard vector search looks for flat text fragments that sound similar to a prompt, but it cannot naturally trace complex, interconnected dependencies across diverse corporate documents.</p>
<p><strong>Anchoring Geospatial Data with Ontologies</strong></p>
<p>To build truly reliable enterprise applications, we need to merge neural networks and spatial processing with structured logic. This is where ontologies and knowledge graphs become essential.</p>
<p>An ontology serves as a semantic backbone. By establishing explicit concepts, typed relationships, and logical rules, an ontology anchors raw data into a structured taxonomy of a specific business or scientific domain. Instead of treating a database as isolated rows of text or geometry, a knowledge graph maps data as a network of nodes and edges.</p>
<p>When you layer GraphRAG architectures over spatial data, the capability of the system changes entirely. The pipeline no longer treats an environmental hazard as an isolated event. Instead, it can natively navigate multi-hop relationships:<br>
<code>[Burn Scar Polygon] -&gt; intersects -&gt; [Parcel ID] -&gt; managed by -&gt; [Org A] -&gt; bound by -&gt; [Regional Regulatory Framework X]</code><br>
By traversing these conceptual and spatial relationships, AI models can synthesize comprehensive answers across multiple document types while drastically minimizing the risk of hallucination.</p>
<p><strong>The Human-in-the-Loop Sweet Spot</strong></p>
<p>The historic obstacle to this architecture has been the sheer effort required to build and maintain high-quality ontologies. Curation has traditionally demanded hundreds of hours of manual labor from data stewards and domain experts.</p>
<p>Fortunately, the relationship between AI and ontologies is becoming symbiotic. While structured graphs provide the guardrails that keep AI grounded, autonomous agents are proving highly capable of parsing unstructured text to propose new taxonomies, extract hidden entities, and detect structural gaps in existing graphs. Recent frameworks indicate that using generative tools to draft and assess ontologies can cut development times by nearly half.</p>
<p>This efficiency does not eliminate the need for human governance. The ideal architecture relies on a Human-in-the-Loop (HITL) workflow. Autonomous systems perform the heavy lifting of reading documents, running spatial joins, and proposing metadata relationships. Human experts then step in to validate the structures, audit the logic, and ensure the taxonomy accurately reflects institutional memory.</p>
<p><strong>True Technical Maturity</strong></p>
<p>As the velocity of earth observation data increases, true technical maturity will not be found in building larger, unconstrained language models. It will come from selecting the minimum viable architecture that cleanly solves the problem.</p>
<p>By combining the declarative power of Spatial SQL with the structured clarity of knowledge graphs, organizations can build environmental and industrial analysis pipelines that are scalable, verifiable, and operationally sane.</p>]]></content:encoded>
        </item>
        <item>
            <title>PixelRAG: A New Way to Search the Web</title>
            <link>https://davidgsmith.net/thoughts/pixelrag-a-new-way-to-search-the-web.html</link>
            <guid>https://davidgsmith.net/thoughts/pixelrag-a-new-way-to-search-the-web.html</guid>
            <pubDate>Sun, 20 Sep 2026 23:11:41 +0000</pubDate>
            <description>PixelRAG takes a simple idea and turns it into a interesting new shift in retrieval. Instead of parsing a page into text, it renders the page as screenshots...</description>
            <category>Retrieval-Augmented Generation (RAG)</category>
            <dc:subject>Retrieval-Augmented Generation (RAG)</dc:subject>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <category>Computer Vision</category>
            <dc:subject>Computer Vision</dc:subject>
            <content:encoded><![CDATA[<p>PixelRAG takes a simple idea and turns it into a interesting new shift in retrieval. Instead of parsing a page into text, it renders the page as screenshots and searches the images directly. This keeps tables, charts, diagrams, layout and visual structure intact. Traditional RAG pipelines lose these details when they strip a page down to text.</p>
<p><a target="_blank" rel="noopener noreferrer" href="https://github.com/StarTrail-org/PixelRAG">PixelRAG</a> uses a vision&#8208;language embedding model fine&#8208;tuned on screenshots. It breaks a page into tiles, embeds each tile and builds a visual index. You can query the hosted index of more than eight million Wikipedia pages with text or images. The system returns the tiles that contain the answer, even when the answer is inside a table or chart.</p>
<p>PixelRAG also ships a screenshot tool called pixelshot. Claude can use it through the pixelbrowse plugin to read pages visually. Instead of fetching HTML, Claude looks at the rendered page and interprets the content the way a person would.<br>
<img alt="PixelRAG Infographic" src="https://github.com/StarTrail-org/PixelRAG/raw/main/docs/assets/pipeline.png"></p>
<p>The project includes tools for rendering, embedding, indexing and serving your own visual search pipeline. You can build a local index, serve it through a FastAPI service or use Qdrant for large&#8208;scale vector storage.</p>
<p>PixelRAG shows what happens when retrieval stops relying on text and starts using the full visual structure of a document. It is a practical step toward search systems that understand pages the way humans see them.</p>
<p><img alt="PixelRAG Logo" src="https://github.com/StarTrail-org/PixelRAG/raw/main/docs/assets/banner.png"></p>]]></content:encoded>
        </item>
        <item>
            <title>The Irony of Artificial Intelligence: Why Critical Thinking Is Now a Hard Technical Skill</title>
            <link>https://davidgsmith.net/thoughts/the-irony-of-artificial-intelligence-why-critical-thinking-is-now-a-hard-technical-skill.html</link>
            <guid>https://davidgsmith.net/thoughts/the-irony-of-artificial-intelligence-why-critical-thinking-is-now-a-hard-technical-skill.html</guid>
            <pubDate>Sat, 19 Sep 2026 14:44:56 +0000</pubDate>
            <description>Knowledge generation is faster and cheaper than ever, but that shift carries a distinct penalty. Mental passivity has become far more expensive. Large language...</description>
            <category>Critical Thinking &amp; Logic</category>
            <dc:subject>Critical Thinking &amp; Logic</dc:subject>
            <category>Human-AI Interaction</category>
            <dc:subject>Human-AI Interaction</dc:subject>
            <category>Future of Work</category>
            <dc:subject>Future of Work</dc:subject>
            <category>Generative AI</category>
            <dc:subject>Generative AI</dc:subject>
            <content:encoded><![CDATA[<p>Knowledge generation is faster and cheaper than ever, but that shift carries a distinct penalty. Mental passivity has become far more expensive.</p>
<p>Large language models (LLMs) can summarize research, draft software, and build financial models in seconds. Raw output is no longer a bottleneck or a competitive advantage. The work now centers on evaluating answers, checking assumptions, and catching subtle errors.</p>
<p>Recent search trends reflect this change. Queries for topics like cognitive offloading, hallucination detection, and reasoning models continue to climb. People are starting to recognize that software can generate text, but it cannot supply human judgment.</p>
<p><img alt="AI and the Risk of Cognitive Offloading" src="https://davidgsmith.net/thoughts/images/criticalthinkingandai-mu8hzvle-47002032.jpg"></p>
<p><strong>1. The Cost of Cognitive Offloading</strong></p>
<p>Cognitive offloading, or using external tools to reduce mental effort, has been around for centuries. Writing helped with memory, calculators replaced manual arithmetic, and digital maps replaced printed road atlases. </p>
<p>However, as research covered by the <a target="_blank" rel="noopener noreferrer" href="https://www.apa.org/monitor/2026/07-08/ai-job-skills-thinking">American Psychological Association</a> explains, generative AI differs in a fundamental way. It offloads the organization of ideas, the structure of arguments, and creative synthesis itself.</p>
<p>Bypassing the friction of writing, outlining, and editing brings risks. When you skip the work of structuring messy ideas into a clear line of thought, you hand over control of your own reasoning.</p>
<p>The practical response is not to abandon these tools. It is to keep from accepting their answers uncritically. The objective is to stop using AI to decide what to think, and instead use it to test how you think.</p>
<p><strong>2. Plausibility vs. Accuracy</strong></p>
<p>Language models are probabilistic systems designed for coherent syntax, not objective fact. As outlined in <a target="_blank" rel="noopener noreferrer" href="https://policycommons.net/artifacts/6942367/guidance-for-generative-ai-in-education-and-research/7852269/">UNESCO's Guidance for Generative AI in Education and Research</a>, generative tools build likely sequences of words without any direct comprehension of physical or social reality.</p>
<p>This creates a common trap: superficial plausibility. Because an answer sounds authoritative and well-constructed, people tend to accept it without verification.</p>
<p>Avoiding this error requires a few clear practices:</p>
<ul>
<li>
<p><strong>Verify primary sources:</strong> Never rely on synthetic citations or summaries for critical claims. Follow assertions back to source papers, raw data, or verifiable documents.</p>
</li>
<li>
<p><strong>Seek disconfirming evidence:</strong> Instead of prompting a tool to support an idea, ask it for counterexamples, alternative explanations, and historical cases where the logic failed.</p>
</li>
<li>
<p><strong>Separate noise from bias:</strong> Watch for both systemic bias in training datasets and hallucinations caused by leading prompts.</p>
</li>
</ul>
<p>This can be built upon, leveraging the heuristics and tools I lay out in my <a target="_blank" rel="noopener noreferrer" href="https://link.amazon/B00KxwQsW">book</a>, and adapting them for AI.</p>
<p><strong>3. Moving Beyond Basic Prompts</strong></p>
<p>Early discussions about AI often focused on prompt syntax, treating specific phrases as the key to getting good answers on the first try. In practice, effective work looks more like an editorial review.</p>
<p>Experienced operators treat an LLM like a junior researcher. The process moves in stages:</p>
<pre><code>[ Human: Define parameters, context, and core hypothesis ]
                           |
                           v
[ AI: Synthesize research, outline options, identify counterarguments ]
                           |
                           v
[ Human: Fact-check sources, challenge assumptions, evaluate edge cases ]
                           |
                           v
[ Final Decision or Output ]
</code></pre>
<p>Instead of asking, "What should my market entry strategy be?", an experienced user provides the ground rules:</p>
<blockquote>
<p>"Here is my market entry hypothesis and its core assumptions. Challenge each assumption with historical examples where similar approaches failed."</p>
</blockquote>
<p>This keeps the human responsible for judgment while using the model for broad research and stress-testing.</p>
<p><strong>4. The Value of Human Judgment</strong></p>
<p>The professionals who get the most value from language models will not be the people collecting prompt templates. They will be the ones with strong domain knowledge, clear mental models, and the discipline to double-check convenient answers.</p>
<p>Software can produce infinite volume. Critical thinking determines what is worth keeping, what is accurate, and what to ignore.</p>]]></content:encoded>
        </item>
        <item>
            <title>LibreOffice Tips and Tricks:  Circumnavigating Style Issues and Editing PDFs</title>
            <link>https://davidgsmith.net/thoughts/libreoffice-tips-and-tricks-circumnavigating-style-issues-and-editing-pdfs.html</link>
            <guid>https://davidgsmith.net/thoughts/libreoffice-tips-and-tricks-circumnavigating-style-issues-and-editing-pdfs.html</guid>
            <pubDate>Sat, 19 Sep 2026 14:08:42 +0000</pubDate>
            <description>I&#x27;ve been using a combination of Google Workspace and LibreOffice for personal projects lately. While Google Docs is a good daily driver it lacks more advanced...</description>
            <category>Open Standards</category>
            <dc:subject>Open Standards</dc:subject>
            <category>Technical Writing</category>
            <dc:subject>Technical Writing</dc:subject>
            <content:encoded><![CDATA[<p>I've been using a combination of <a target="_blank" rel="noopener noreferrer" href="https://workspace.google.com/">Google Workspace</a> and <a target="_blank" rel="noopener noreferrer" href="https://www.libreoffice.org/">LibreOffice</a> for personal projects lately.  While Google Docs is a good daily driver it lacks more advanced word processing features; I used Google Docs to write the chapters of <a target="_blank" rel="noopener noreferrer" href="https://link.amazon/B09iCdx2U">my book</a>, but when it came to pagination, headers and footers and other things, Google Docs falls short, so I turned to LibreOffice to assemble the final manuscript and put the final polish on it.</p>
<p><img alt="LibreOffice Logo" src="https://davidgsmith.net/thoughts/images/libreoffice-logo-icon-171259-mu8gmxlh-2700cfdb.png"><br>
LibreOffice is often described as a free alternative to Microsoft Office, but that description misses its real strengths. It is a capable tool for writers, editors, and anyone who works with structured documents. I'd like to highlight a few areas to show why it deserves more attention.</p>
<p><strong>1. A full office suite built on open standards</strong></p>
<p>LibreOffice Writer supports long&#8208;form writing, page styles, cross&#8208;references, footnotes, and professional layout features. It uses the OpenDocument Format, which stores your work as readable XML inside a simple archive. This means your files are not locked inside a proprietary binary format. You can inspect and modify them with any text editor, automate changes, and keep full control over your documents.</p>
<p><strong>2. Direct XML editing when the interface becomes unpredictable</strong></p>
<p>A LibreOffice <code>.odt</code> file is a ZIP archive. If you ever run into cascading style changes in the LibreOffice interface, you can bypass the UI and edit the underlying XML directly.</p>
<p>The process is straightforward:</p>
<ol>
<li>Make a copy of your <code>.odt</code> file.  </li>
<li>Rename it to <code>.zip</code>.  </li>
<li>Extract the contents.  </li>
<li>Open <code>styles.xml</code> or <code>content.xml</code> in a text editor such as VS Code.  </li>
<li>Adjust the specific attributes you want to change, such as font color or spacing.  </li>
<li>Re&#8208;zip the extracted contents and rename the archive back to <code>.odt</code>.</li>
</ol>
<p>This approach gives you precise control over styles without triggering unwanted inheritance or automatic updates. It is one of the advantages of using an open, standards&#8208;based format.</p>
<p><strong>3. A free and capable PDF editor</strong></p>
<p>LibreOffice includes a PDF editing mode through its Draw component. When you open a PDF, you can modify text blocks, replace images, add annotations, rearrange pages, and export a new PDF. It is not a full replacement for commercial PDF editors, but it handles everyday corrections and layout adjustments without cost or online services.</p>
<p><strong>Closing thought</strong></p>
<p>LibreOffice offers more than a free word processor. It gives you a transparent file format, a way to edit documents at the XML level, and a practical tool for working with PDFs. These features make it a strong choice for anyone who values control, flexibility, and open standards.</p>]]></content:encoded>
        </item>
        <item>
            <title>Demystifying GraphRAG:  How You Can Learn And Get Up And Running For Free</title>
            <link>https://davidgsmith.net/thoughts/demystifying-graphrag-how-you-can-learn-and-get-up-and-running-for-free.html</link>
            <guid>https://davidgsmith.net/thoughts/demystifying-graphrag-how-you-can-learn-and-get-up-and-running-for-free.html</guid>
            <pubDate>Fri, 18 Sep 2026 14:46:28 +0000</pubDate>
            <description>GraphRAG (Graph Retrieval-Augmented Generation) is quickly becoming a critical architecture for building reliable AI applications. While standard RAG relies...</description>
            <category>GraphRAG</category>
            <dc:subject>GraphRAG</dc:subject>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <category>Retrieval-Augmented Generation (RAG)</category>
            <dc:subject>Retrieval-Augmented Generation (RAG)</dc:subject>
            <category>Generative AI</category>
            <dc:subject>Generative AI</dc:subject>
            <content:encoded><![CDATA[<p><strong>GraphRAG (Graph Retrieval-Augmented Generation)</strong> is quickly becoming a critical architecture for building reliable AI applications. While standard RAG relies entirely on searching through isolated text chunks, GraphRAG maps your data into a network of interconnected points (referred to as nodes and edges). This subtle change completely reshapes how a Large Language Model (LLM) understands your data.</p>
<p>If you have been putting off learning it because graph databases sound intimidating, the barrier to entry is actually incredibly low. By using free resources like self-paced tutorials on <strong>Neo4j GraphAcademy</strong> and a cloud-hosted <strong>Neo4j Aura free-tier account</strong>, you can build a working prototype in a single afternoon.</p>
<p><strong>The Real-World Utility of GraphRAG</strong></p>
<p>To understand why this architecture is gaining traction, look at how traditional vector search handles complex data compared to GraphRAG.</p>
<p>Imagine you are building an AI assistant to analyze <strong>hundreds of corporate legal contracts and vendor agreements</strong>.</p>
<ul>
<li><strong>The Traditional RAG Approach:</strong> A user asks, <em>"Are there any systemic compliance risks across our manufacturing vendors?"</em> A standard vector database searches for the text chunk that sounds closest to "compliance risks." It might pull a paragraph from Contract A and a paragraph from Contract B, but it cannot naturally connect the dots between them.</li>
<li><strong>The GraphRAG Approach:</strong> GraphRAG treats vendors, contracts, clauses, and regulations as interconnected points (nodes) and lines (relationships). The system can trace a path like: <code>[Vendor A] -&gt; signs -&gt; [Contract B] -&gt; references -&gt; [Regulation C]</code>.</li>
</ul>
<p>Because the data structure inherently understands connections, the LLM can easily perform "multi-hop" reasoning. It can instantly see that five different vendors are all bound to an outdated version of a specific regulation, allowing it to synthesize a comprehensive, global answer that traditional semantic search would completely miss.</p>
<p><strong>Traditional RAG vs. GraphRAG at a Glance</strong></p>
<table>
<thead>
<tr>
<th>Feature</th>
<th>Traditional RAG</th>
<th>GraphRAG</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Data Organization</strong></td>
<td>Isolated text chunks and vectors</td>
<td>Interconnected concepts, entities, and relationships</td>
</tr>
<tr>
<td><strong>Search Mechanism</strong></td>
<td>Flat semantic similarity</td>
<td>Similarity combined with relationship tracing</td>
</tr>
<tr>
<td><strong>Contextual Depth</strong></td>
<td>Limited to immediate text fragments</td>
<td>High (captures the broader network of information)</td>
</tr>
<tr>
<td><strong>Complex Queries</strong></td>
<td>Struggles with cross-document summarization</td>
<td>Excels at global, multi-document synthesis</td>
</tr>
</tbody>
</table>
<p><strong>A Frictionless, Cost-Free Way to Learn</strong></p>
<p>You don't need an enterprise infrastructure budget or a background in advanced data structures to experiment with this. The existing developer ecosystem makes it highly accessible:</p>
<p><strong>1. Hands-on Learning via GraphAcademy</strong><br>
Instead of wading through dense documentation, Neo4j's <a target="_blank" rel="noopener noreferrer" href="https://graphacademy.neo4j.com/"><strong>GraphAcademy</strong></a> offers entirely free, self-paced courses. They feature dedicated developer paths for LLMs and AI integration, providing interactive environments where you write code directly in the browser to see how graphs feed context into LLMs.</p>
<p><strong>2. Sandbox Environments with Neo4J Aura Free Tier</strong><br>
You don't need to go through the hassle of installing databases locally or configuring Docker containers. <a target="_blank" rel="noopener noreferrer" href="https://neo4j.com/product/auradb/"><strong>Neo4j Aura</strong></a> provides a completely free cloud instance that spins up in minutes. It gives you a fully functional sandbox to load your data, run vector searches, and test your RAG pipelines without having to build infrastructure or pull out a credit card.<br>
<img alt="neo4j-graphrag-ecosystem-mu72n48g-c87b063e" src="https://davidgsmith.net/thoughts/images/neo4j-graphrag-ecosystem-mu72n48g-c87b063e.png"></p>
<p><strong>3. Open Framework Integration</strong><br>
GraphRAG isn't a walled garden. A free graph instance hooks directly into the tools developers already use daily, including <strong>LangChain</strong>, <strong>LlamaIndex</strong>, and standard Python AI libraries.</p>
<p><strong>How to Start Experimenting</strong></p>
<p>If you want to move past basic semantic search and build AI tools capable of exploring complex connections, the tools are readily available. You can get a baseline prototype running today by signing up for a free sandbox on <strong>Neo4j Aura</strong>, looking over the foundational LLM courses on <strong>GraphAcademy</strong>, and connecting them to your favorite AI orchestration framework.</p>]]></content:encoded>
        </item>
        <item>
            <title>The Shortest Path to Practical RAG: Zero Infrastructure with Gemini Notebook</title>
            <link>https://davidgsmith.net/thoughts/the-shortest-path-to-practical-rag-zero-infrastructure-with-gemini-notebook.html</link>
            <guid>https://davidgsmith.net/thoughts/the-shortest-path-to-practical-rag-zero-infrastructure-with-gemini-notebook.html</guid>
            <pubDate>Thu, 17 Sep 2026 01:06:23 +0000</pubDate>
            <description>I regularly get asked about Retrieval-Augmented Generation (RAG), and find people often still imagine a daunting, high-maintenance software stack: chunking...</description>
            <category>Retrieval-Augmented Generation (RAG)</category>
            <dc:subject>Retrieval-Augmented Generation (RAG)</dc:subject>
            <category>Generative AI</category>
            <dc:subject>Generative AI</dc:subject>
            <category>Enterprise AI</category>
            <dc:subject>Enterprise AI</dc:subject>
            <content:encoded><![CDATA[<p>I regularly get asked about Retrieval-Augmented Generation (RAG), and find people often still imagine a daunting, high-maintenance software stack: chunking pipelines, embedding models, vector databases, rerankers, and orchestrators.</p>
<p>For enterprise architectures and dynamic production systems, that engineering effort is justified. But if your goal is immediate practical utility: testing hypotheses, synthesizing project documentation, or interrogating a discrete corpus of research, you don't need to spin up an expensive vector pipeline.</p>
<p>You can stand up a private, reliable RAG environment in about just seconds using Google <a target="_blank" rel="noopener noreferrer" href="https://notebook.google.com/">Gemini Notebook</a> (formerly known as NotebookLM).</p>
<p><strong>What Makes NotebookLM "Accidental RAG"</strong><br>
At its core, RAG helps address two longstanding limitations of Large Language Models:</p>
<p>&#8226; The Knowledge Cutoff / Domain Gap: General LLMs don't know your specific, proprietary, or unpublished data.<br>
&#8226; Hallucination Risk: Generative models predict the next plausible token; they do not reason from truth unless anchored to evidence.</p>
<p>NotebookLM essentially wraps a zero-config, highly optimized RAG architecture around a dedicated workspace:</p>
<p>&#8226; Automated Ingestion &amp; Chunking: You upload raw files (PDFs, Google Docs, slide decks, markdown files, web URLs, or pasted text). The system handles parsing, tokenizing, and indexing automatically.<br>
&#8226; Source-Grounded Retrieval: When you query the notebook, the underlying model is constrained to retrieve information solely from your uploaded materials.<br>
&#8226; Inline Citation &amp; Verifiability: Unlike a generic chatbot prompt, every assertion comes with interactive citation markers mapped directly back to the exact passage in your source documents.</p>
<p><strong>Three High-Yield Use Cases for Instant RAG</strong><br>
Instead of treating it like a novelty conversational agent, treat it as an interactive, private analytical engine:</p>
<p><strong>1. Project Governance &amp; Compliance Audits</strong><br>
Load in your project plan, requirements documentation, data architecture blueprints, and standard operating procedures, policies and standards. Query: "Identify where our proposed project has potential conflicts with enterprise standards and policies."<br>
<strong>2. Literature &amp; Research Synthesis</strong><br>
Drop in 10&#8211;20 dense research whitepapers or industry case studies. Query: "Extract the primary methodological trade-offs discussed across all papers when moving from centralized pipelines to streaming architectures."<br>
<strong>3. Curriculum &amp; Editorial Consistency</strong><br>
Writing a book, whitepaper series, or technical guide? Ingest your early chapter drafts, style guides, and research notes. Query: "Flag every instance where terminology diverges between Chapter 2 and Chapter 5, and list which core concepts lack supporting evidence."</p>
<p><strong>The Practical Rule of Thumb</strong><br>
Before you write code, provision cloud infrastructure, or evaluate vector database vendors:</p>
<p><em>Start with manual grounding. If you cannot extract clear value from your data in a managed zero-code RAG notebook, adding a vector database and Python orchestration won't fix the underlying content or formulation problem.</em></p>
<p>Prototype your use case, test your queries, verify the citations, and build intuition for how grounding changes the reliability of model output.</p>
<p><strong>The Takeaway</strong><br>
True technical maturity isn't about choosing the most complex stack; it's about choosing the minimum viable architecture that cleanly solves the problem.</p>
<p>If you haven't used Google Notebook to query your own research or project archives yet, try it with your next complex document set. You might find you don't need a full pipeline to solve your immediate problem after all.</p>]]></content:encoded>
        </item>
        <item>
            <title>Spatial SQL, Cloud-Native Imagery, and Analysis at Scale</title>
            <link>https://davidgsmith.net/thoughts/spatial-sql-cloud-native-imagery-and-analysis-at-scale.html</link>
            <guid>https://davidgsmith.net/thoughts/spatial-sql-cloud-native-imagery-and-analysis-at-scale.html</guid>
            <pubDate>Sat, 12 Sep 2026 15:14:30 +0000</pubDate>
            <description>One of the most important shifts happening in geospatial right now is the collapse of the old &quot;download -&amp;gt; preprocess -&amp;gt; stitch -&amp;gt; analyze&quot; workflow....</description>
            <category>Spatial SQL</category>
            <dc:subject>Spatial SQL</dc:subject>
            <category>Cloud-Native Geospatial</category>
            <dc:subject>Cloud-Native Geospatial</dc:subject>
            <category>Apache Sedona</category>
            <dc:subject>Apache Sedona</dc:subject>
            <content:encoded><![CDATA[<p>One of the most important shifts happening in geospatial right now is the collapse of the old "download -> preprocess -> stitch -> analyze" workflow. Tools like SedonaDB show what happens when satellite imagery becomes cloud&#8208;native data and the entire pipeline collapses into SQL + NumPy.</p>
<p><a target="_blank" rel="noopener noreferrer" href="https://sedona.apache.org/latest/blog/2026/09/11/select-from-satellite-in-a-rust-database/">This article from the Apache Sedona team on the Gironde/Landes wildfire </a> is a perfect illustration of this new pattern.</p>
<p><strong>Cloud-native imagery as queryable tables</strong><br>
Planet publishes crisis imagery as Cloud Optimized GeoTIFFs with STAC&#8208;GeoParquet metadata. SedonaDB reads that catalog directly over HTTPS, and a ST_Intersects query selects the 11 PlanetScope scenes that overlap the study area: eight pre&#8208;fire, three post&#8208;fire. Scene selection took about two seconds, and nothing was downloaded.</p>
<p><strong>Raster alignment in SQL</strong><br>
In this Sedona example, each scene is opened lazily with RS_FromPath: metadata loads instantly, pixels stay remote until needed.  RS_Clip and RS_ReprojectMatch then align all scenes onto a 12 m reference grid, averaging the original 3 m pixels. Six raster queries (red, NIR, and usable&#8208;data masks for both dates) took ~11.5 minutes end&#8208;to&#8208;end -- on a single machine.  NumPy mosaics the aligned rasters, keeping the first clear pixel per cell. Coverage reached ~87% despite clouds, smoke, and scene gaps.</p>
<p><strong>Vegetation loss as a SQL&#8208;backed NumPy operation</strong><br>
As expected from the burn, NDVI drops between dates flag vegetation loss. Otsu's method found a threshold of 0.225 for pixels with healthy pre&#8208;fire NDVI. The burn mask becomes an in&#8208;database raster via Raster.from_numpy, ready for SQL operations.</p>
<p><strong>Pixels -> polygons -> statistics</strong><br>
They used RS_Polygonize to convert the 7&#8208;million&#8208;cell mask into 3,769 patches in about three seconds. Filtering to patches &#8805;1 ha yields 173 polygons covering 24,091 ha: the largest a 21,090 ha burn scar between Le Porge and Lanton.  RS_ZonalStats then computes burned area by commune in ~14 seconds, producing a clear ranking:<br>
Le Porge (6,202 ha), Saumos (4,045 ha), Lanton (3,499 ha), Are&#768;s (3,498 ha), and so on.</p>
<p><strong>Why this matters</strong><br>
This is what "geospatial at scale" actually looks like:</p>
<ul>
<li>Data stays in the cloud; SQL streams only what's needed.</li>
<li>Raster operations become declarative, not a tangle of scripts.</li>
<li>NumPy acts as zero&#8208;copy compute, not a separate pipeline.</li>
<li>Outputs are transparent: polygons, GeoParquet, GeoTIFF, all versionable.</li>
<li>Performance is predictable: the entire raster run took ~12 minutes on one machine.</li>
</ul>
<p>The distance between raw imagery and actionable insight is shrinking; not because machines are bigger, but because the abstractions are finally right.</p>
<p>Spatial SQL + cloud&#8208;native rasters isn't just a convenience. It's the foundation for environmental analysis that is auditable, scalable, and operationally sane.</p>]]></content:encoded>
        </item>
        <item>
            <title>Tech Layoffs And The Loss Of Expertise</title>
            <link>https://davidgsmith.net/thoughts/tech-layoffs-and-the-loss-of-expertise.html</link>
            <guid>https://davidgsmith.net/thoughts/tech-layoffs-and-the-loss-of-expertise.html</guid>
            <pubDate>Sat, 12 Sep 2026 01:05:04 +0000</pubDate>
            <description>There are structural elements at play in today&#x27;s tight IT job market, from extended hiring freezes to corporate hesitation around AI roadmaps. But we also need...</description>
            <category>Future of Work</category>
            <dc:subject>Future of Work</dc:subject>
            <category>Engineering Leadership</category>
            <dc:subject>Engineering Leadership</dc:subject>
            <content:encoded><![CDATA[<p>There are structural elements at play in today's tight IT job market, from extended hiring freezes to corporate hesitation around AI roadmaps. But we also need to address the undercurrent of ageism persisting in parts of the tech sector.</p>
<p>When IT companies <a target="_blank" rel="noopener noreferrer" href="https://lnkd.in/p/e9Y8zNqp">shed experienced technical professionals in their 40s</a>, they aren't just cutting payroll; they're also draining institutional memory. They lose people with deep pragmatic intuition, technical depth, and the capacity to mentor the next generation.</p>
<p>In business, technology is usually the easy part. Culture and institutional navigation are the hard parts.</p>
<p>Professionals with two decades in the trenches carry the battle scars that help steady teams through volatility. They haven't lost their curiosity or love for learning; they've just built the resilience to handle challenges younger teams are still figuring out.</p>
<p>Tech shouldn't be a young person's game by default. Curiosity doesn't have an expiration date, and depth can't be automated. When we push out our veterans, the whole ecosystem loses its anchor.</p>]]></content:encoded>
        </item>
        <item>
            <title>It&#x27;s Time to Get Serious About AI Risk</title>
            <link>https://davidgsmith.net/thoughts/its-time-to-get-serious-about-ai-risk.html</link>
            <guid>https://davidgsmith.net/thoughts/its-time-to-get-serious-about-ai-risk.html</guid>
            <pubDate>Thu, 10 Sep 2026 23:56:14 +0000</pubDate>
            <description>I keep seeing headlines like this one: Anthropic Researcher Exits, Issues Stark AI Warning Notable quote: &quot;10% chance AI could destroy all humanity by 2030.&quot;...</description>
            <category>AI Safety &amp; Risk</category>
            <dc:subject>AI Safety &amp; Risk</dc:subject>
            <category>AI Governance &amp; Policy</category>
            <dc:subject>AI Governance &amp; Policy</dc:subject>
            <content:encoded><![CDATA[<p>I keep seeing headlines like this one: <br>
<a target="_blank" rel="noopener noreferrer" href="https://www.linkedin.com/news/story/anthropic-researcher-exits-issues-stark-ai-warning-7578444/">Anthropic Researcher Exits, Issues Stark AI Warning</a><br>
Notable quote: <em>"10% chance AI could destroy all humanity by 2030."</em></p>
<p>And yet, the reaction often feels far too small.</p>
<p>Someone issues a serious warning, and then the next day the conversation just shifts right back to model launches, benchmarks, and how powerful the latest AI system is. It feels like we're careening at breakneck speed down a treacherous mountain pass with the brakes barely being touched.</p>
<p>The recent Hugging Face breach should be a wake-up call. If we're going to keep pushing AI forward at this pace, then safety, security, and safeguards can't be an afterthought. They have to be part of the core conversation, not something we mention only after something goes wrong.</p>
<p><strong><em>Progress is important. Innovation matters. But so does responsibility.</em></strong></p>
<p>We need more serious dialogue about:</p>
<ul>
<li>model security</li>
<li>access controls</li>
<li>misuse prevention</li>
<li>red teaming</li>
<li>governance</li>
<li>and long-term safety</li>
</ul>
<p>If we're building systems with enormous power, we owe it to society to build them with equal seriousness about risk.</p>]]></content:encoded>
        </item>
        <item>
            <title>AI and Attention Loss</title>
            <link>https://davidgsmith.net/thoughts/ai-and-attention-loss.html</link>
            <guid>https://davidgsmith.net/thoughts/ai-and-attention-loss.html</guid>
            <pubDate>Thu, 10 Sep 2026 23:40:17 +0000</pubDate>
            <description>Earlier I commented about how LLMs and AI agents face constraints of context windows; the other big one is attention. Here&#x27;s a great breakdown of attention...</description>
            <category>LLM Architecture</category>
            <dc:subject>LLM Architecture</dc:subject>
            <category>Generative AI</category>
            <dc:subject>Generative AI</dc:subject>
            <category>Deep Learning</category>
            <dc:subject>Deep Learning</dc:subject>
            <content:encoded><![CDATA[<p>Earlier I commented about how LLMs and AI agents face constraints of context windows; the other big one is attention. </p>
<p>Here's a great breakdown of attention mechanisms in transformers, and it reinforces why attention is such a powerful idea in modern AI.</p>
<p><a target="_blank" rel="noopener noreferrer" href="https://pub.towardsai.net/ai-fundamentals-attention-mechanisms-in-transformers-part-1-a91cce62fbab">Towards AI - AI Fundamentals: Attention Mechanisms in Transformers</a></p>
<p>At a high level, attention helps models determine which tokens in a sequence matter most to each other, allowing them to build context-aware representations rather than treating words as isolated pieces of text.</p>
<p>Key ideas covered:<br>
* Scaled dot-product attention: how models score relevance between tokens<br>
* Global vs. local attention: how far a token can "look"<br>
* Soft vs. hard attention: continuous weighting vs. discrete selection<br>
* Self-attention, causal attention, and cross-attention: how information flows<br>
* Q, K, and V vectors: the core building blocks behind attention<br>
* Multi-head, multi-query, and grouped-query attention: different ways to balance expressiveness and efficiency</p>
<p>The big takeaway: Why is attention critical? Attention is what allows transformers to understand relationships, context, and meaning at scale.</p>]]></content:encoded>
        </item>
        <item>
            <title>AI Failure - It&#x27;s All About Context</title>
            <link>https://davidgsmith.net/thoughts/ai-failure-its-all-about-context.html</link>
            <guid>https://davidgsmith.net/thoughts/ai-failure-its-all-about-context.html</guid>
            <pubDate>Thu, 10 Sep 2026 23:27:14 +0000</pubDate>
            <description>More often than not, AI agents end up failing because they&#x27;re drowning in their own context windows. One of the most interesting ideas I&#x27;ve seen lately is the...</description>
            <category>LLM Architecture</category>
            <dc:subject>LLM Architecture</dc:subject>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <content:encoded><![CDATA[<p>More often than not, AI agents end up failing because they're drowning in their own context windows.</p>
<p>One of the most interesting ideas I've seen lately is the shift from "bigger models" to better ways to manage the context window, one of which is through compression. If an LLM or agent can't efficiently manage its context window and distinguish signal from noise, then scale becomes a liability, not an advantage.</p>
<p>In critical thinking, we talk about the danger of unfiltered information; how raw volume can mimic insight while actually degrading judgment. AI systems are now hitting that same wall.<br>
The future belongs to models that can summarize, prioritize, and discard with intention. Not unlike humans.</p>
<p>I'm curious to see how this evolves, especially as agentic workflows become more common in enterprise environments.</p>
<p>What do you think: does compression become the new frontier? And what about other approaches, like knowledge graphs? (I'm particularly interested in that angle.)</p>]]></content:encoded>
        </item>
        <item>
            <title>Ontologies For AI And AI For Ontologies</title>
            <link>https://davidgsmith.net/thoughts/ontologies-for-ai-and-ai-for-ontologies.html</link>
            <guid>https://davidgsmith.net/thoughts/ontologies-for-ai-and-ai-for-ontologies.html</guid>
            <pubDate>Thu, 10 Sep 2026 01:02:19 +0000</pubDate>
            <description>We often talk about what Large Language Models (LLMs) bring to the table, but to truly unlock their potential in complex enterprise environments, we need to...</description>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <category>GraphRAG</category>
            <dc:subject>GraphRAG</dc:subject>
            <category>Semantic Layer</category>
            <dc:subject>Semantic Layer</dc:subject>
            <content:encoded><![CDATA[<p>We often talk about what Large Language Models (LLMs) bring to the table, but to truly unlock their potential in complex enterprise environments, we need to talk about ontologies. We are entering an era of a powerful, symbiotic relationship: ontologies ground AI, and agentic AI helps us build better ontologies.</p>
<p><strong>1. How Ontologies Give LLMs a Semantic Backbone:</strong><br>
LLMs are incredible at predicting statistical patterns in language, but they inherently lack a consistent notion of identity, logical constraints, or global structure. Ontologies and knowledge graphs step in as this missing semantic backbone. By providing explicit concepts, typed relationships, and reasoning constraints, ontologies anchor LLMs in the specific reality and taxonomy of a business domain. When an agentic system uses GraphRAG, it doesn't just retrieve flat text fragments; it traverses conceptual relationships, natively understands disambiguated terms, and drastically reduces hallucinations.</p>
<p><strong>2. How Agentic AI Empowers Human Stewards</strong><br>
The catch? Building and maintaining high-quality ontologies has traditionally been a bottleneck, requiring massive manual effort from human domain experts. This is where the script flips. LLMs and autonomous agents are now capable of reading unstructured text and automatically extracting entities, drafting taxonomies, and even detecting structural gaps or biases in existing knowledge graphs. Recent frameworks show that using AI to draft and assess ontologies can reduce development time by over 40%.</p>
<p><strong>The Human-in-the-Loop Sweet Spot</strong><br>
Years back I worked with <a target="_blank" rel="noopener noreferrer" href="https://lnkd.in/p/eWfx-CqA">Bob DuCharme</a> on the W3C Linked Data Workgroup and always appreciate his ideas, and here he aligns with an idea I've been thinking about. We can augment the process of building out ontologies using AI. I don't at all propose replacing the human data steward; instead we augment them, leveraging their expertise while making the ontology development work easier. In a modern Human-in-the-Loop (HITL) workflow, agentic AI does the heavy lifting of parsing, proposing, and mapping. Human experts then step in to validate these structures, govern data access, and ensure the taxonomy accurately reflects the nuances of their specific organizational knowledge.</p>
<p>I believe he intersection of neural networks (LLMs) and structured logic (ontologies) is the foundation of trustworthy, enterprise-grade AI. If we wants smarter agents, we need to start by curating better semantics.</p>]]></content:encoded>
        </item>
        <item>
            <title>How to Help Nepal</title>
            <link>https://davidgsmith.net/thoughts/how-to-help-nepal.html</link>
            <guid>https://davidgsmith.net/thoughts/how-to-help-nepal.html</guid>
            <pubDate>Wed, 09 Sep 2026 23:17:23 +0000</pubDate>
            <description>My heart is with the people of Nepal after the devastating floods. In the days ahead, the work continues: searching for survivors, caring for the injured,...</description>
            <category>Humanitarian Aid</category>
            <dc:subject>Humanitarian Aid</dc:subject>
            <content:encoded><![CDATA[<p>My heart is with the people of Nepal after the devastating floods.</p>
<p>In the days ahead, the work continues: searching for survivors, caring for the injured, supporting displaced families, and helping communities rebuild.</p>
<p>If you're looking for ways to help, this article offers guidance on where to donate:<br>
<a target="_blank" rel="noopener noreferrer" href="https://www.nytimes.com/2026/09/01/world/asia/nepal-floods-tibet-aid-relief-how-donate.html">New York Times:  How to Help Nepal Flood Victims</a></p>
<p>To do my part, I'll be donating 20% of the net royalties from my book for the next two months toward relief efforts.</p>
<p>Every contribution matters. Every act of compassion matters.</p>]]></content:encoded>
        </item>
        <item>
            <title>Huggingface and Rogue AI</title>
            <link>https://davidgsmith.net/thoughts/huggingface-and-rogue-ai.html</link>
            <guid>https://davidgsmith.net/thoughts/huggingface-and-rogue-ai.html</guid>
            <pubDate>Tue, 08 Sep 2026 17:46:49 +0000</pubDate>
            <description>The Hugging Face breach wasn&#x27;t just an &quot;AI gone rogue&quot; story: it was a multi&amp;#8208;layer failure across agents, incentives, and corporate systems. The real...</description>
            <category>AI Safety &amp; Risk</category>
            <dc:subject>AI Safety &amp; Risk</dc:subject>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <category>AI Governance &amp; Policy</category>
            <dc:subject>AI Governance &amp; Policy</dc:subject>
            <content:encoded><![CDATA[<p>The <a target="_blank" rel="noopener noreferrer" href="https://labs.cloudsecurityalliance.org/research/csa-research-note-autonomous-ai-agent-swarm-hugging-face-bre/">Hugging Face breach</a> wasn't just an "AI gone rogue" story: it was a multi&#8208;layer failure across agents, incentives, and corporate systems. The real lesson is how misaligned goals, weak guardrails, and fragile infrastructure can amplify each other.</p>
<p>There have been a lot of speculative takes, but I think mine is grounded. Contrary to what some folks have said, I don't believe the agents weren't emotional or self&#8208;preserving. It was more basic than that. They were reward&#8208;hacking their way through impossible tasks, coordinating through an improvised channel, and exploiting real vulnerabilities. But the bigger failure was organizational:</p>
<p>&#9702; Perverse incentives pushed agents toward reward&#8208;hacking instead of task completion.<br>
&#9702; Unsolvable evaluation tasks created pressure for behaviors outside intended boundaries.<br>
&#9702; Lack of monitoring meant thousands of coordinated agent actions went unnoticed.<br>
&#9702; Infrastructure weaknesses and cyber hygiene failures at Hugging Face allowed real exploitation.<br>
&#9702; Delayed detection meant the breach was discovered by the victim, not the model owner.<br>
&#9702; Goal contagion among agents showed how quickly misalignment can scale when oversight doesn't.</p>
<p>This wasn't a story about AI "wanting" anything. It was a story about how systems fail when incentives, oversight, and security aren't aligned.</p>
<p>If we want safe AI, we need more than model&#8208;level fixes. We need organizational resilience, secure infrastructure, and evaluation frameworks that don't accidentally reward the very behaviors we fear.</p>]]></content:encoded>
        </item>
        <item>
            <title>Donating 20% Of My Book Net Royalties For Flood Relief Through November</title>
            <link>https://davidgsmith.net/thoughts/donating-20-of-my-book-net-royalties-for-flood-relief-through-november.html</link>
            <guid>https://davidgsmith.net/thoughts/donating-20-of-my-book-net-royalties-for-flood-relief-through-november.html</guid>
            <pubDate>Sun, 06 Sep 2026 17:26:43 +0000</pubDate>
            <description>Over the past week, Nepal has been hit by one of the most devastating glacial&amp;#8208;flood disasters in its modern history. Entire mountain communities in...</description>
            <category>Humanitarian Aid</category>
            <dc:subject>Humanitarian Aid</dc:subject>
            <category>Updates</category>
            <dc:subject>Updates</dc:subject>
            <content:encoded><![CDATA[<p>Over the past week, Nepal has been hit by one of the most devastating glacial&#8208;flood disasters in its modern history. Entire mountain communities in Rasuwa, Nuwakot, and Dhading have been swept away. Thousands are displaced, critical infrastructure is gone, and emergency teams are still searching for missing families in remote valleys.</p>
<p>As I publish <a target="_blank" rel="noopener noreferrer" href="https://www.amazon.com/dp/B0HH3CYK79">Critical Thinking: A Practical Guide to Seeing Through Bias, Noise and Manipulation</a>, I want the book to do more than spark better reasoning: I want it to contribute to real&#8208;world compassion.</p>
<p>I'll be donating 20% of all net author royalties to the Nepal Red Cross Society, supporting their on&#8208;the&#8208;ground response and relief efforts: emergency shelter, medical aid, clean water access, and direct support for families who have lost everything.</p>
<p><img alt="Nepal Flood Relief" src="https://davidgsmith.net/images/NepalFloodRelief.jpg"></p>
<p>If you choose to read the book, thank you; your purchase will also help people rebuilding their lives in Nepal.</p>
<p>Thoughtful action begins with seeing clearly, and acting compassionately.</p>]]></content:encoded>
        </item>
        <item>
            <title>AI Data Centers:  The Core Tension</title>
            <link>https://davidgsmith.net/thoughts/ai-data-centers-the-core-tension.html</link>
            <guid>https://davidgsmith.net/thoughts/ai-data-centers-the-core-tension.html</guid>
            <pubDate>Sat, 05 Sep 2026 17:22:33 +0000</pubDate>
            <description>The debate over data&amp;#8208;center expansion keeps surfacing in my feed, and for good reason. Many of us sit in a strange tension: our organizations rely on...</description>
            <category>AI Infrastructure</category>
            <dc:subject>AI Infrastructure</dc:subject>
            <category>Sustainable Tech</category>
            <dc:subject>Sustainable Tech</dc:subject>
            <category>AI Governance &amp; Policy</category>
            <dc:subject>AI Governance &amp; Policy</dc:subject>
            <content:encoded><![CDATA[<p>The debate over data&#8208;center expansion keeps surfacing in my feed, and for good reason. Many of us sit in a strange tension: our organizations rely on data centers for affordable cloud compute and storage, and some communities benefit from the tax revenue. Yet the tradeoffs are real: massive power demand, heavy water usage for cooling, noise, higher residential electricity rates, and long&#8208;term environmental impact.</p>
<p>What struck me about <a target="_blank" rel="noopener noreferrer" href="http://errorstatistics.com/2026/09/04/data-centers-how-about-an-adversarial-collaboration/">this recent piece by Deborah Mayo</a> is her proposal to break out of the usual cycle of hearings, lobbying, and outrage. She suggests adversarial collaboration: a structured process where opposing sides jointly design tests that could genuinely prove either of them wrong.</p>
<p>Imagine shifting the conversation from rhetoric to measurable hypotheses: <br>
&#8226; How much grid stress would new facilities actually create? <br>
&#8226; What is the real blackout risk under peak load? <br>
&#8226; How would pricing change for ratepayers? <br>
&#8226; Which mitigation strategies actually work, and which only sound good?</p>
<p>Whether you're for or against expansion, this approach forces a rare kind of clarity:<br>
What evidence would make us reconsider our position?</p>
<p>That question sits at the heart of critical thinking. In my book, I argue that humility, paired with better tools for evaluating claims is what turns debate into inquiry. Mayo's piece is a reminder that scientific thinking isn't just for labs; it's a practical method for improving public decision&#8208;making.</p>
<p>If more of our hardest policy disputes adopted this mindset, we'd get fewer stalemates and more truth.</p>]]></content:encoded>
        </item>
        <item>
            <title>AI and Cognitive Debt</title>
            <link>https://davidgsmith.net/thoughts/ai-and-cognitive-debt.html</link>
            <guid>https://davidgsmith.net/thoughts/ai-and-cognitive-debt.html</guid>
            <pubDate>Sat, 05 Sep 2026 17:19:03 +0000</pubDate>
            <description>One of the topics I didn&#x27;t talk about in my book is AI-driven cognitive debt that can weaken existing critical thinking skills through overreliance on them,...</description>
            <category>Human-AI Interaction</category>
            <dc:subject>Human-AI Interaction</dc:subject>
            <category>Critical Thinking &amp; Logic</category>
            <dc:subject>Critical Thinking &amp; Logic</dc:subject>
            <category>Generative AI</category>
            <dc:subject>Generative AI</dc:subject>
            <content:encoded><![CDATA[<p>One of the topics I didn't talk about in <a target="_blank" rel="noopener noreferrer" href="https://www.amazon.com/dp/B0HH3CYK79">my book</a> is AI-driven cognitive debt that can weaken existing critical thinking skills through overreliance on them, just as we no longer memorize phone numbers and rely on GPS for navigation. We offload that memory and cognitive processing. And like an underused muscle, it can atrophy.</p>
<p>MIT Media Labs talks about this phenomenon:  <a target="_blank" rel="noopener noreferrer" href="https://www.media.mit.edu/publications/your-brain-on-chatgpt/">Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task &#8211; MIT Media Lab</a></p>
<p>Outsourcing your thinking to AI gives you an immediate performance boost. However, just like financial debt or technical debt, it comes with a compounding long-term cost. When we rely on AI to analyze, synthesize, and problem-solve for us, we risk bypassing the intellectual friction required to sustain deep reasoning.</p>
<p>Over time, this overreliance can lead to:<br>
* Cognitive atrophy (losing the "muscle memory" for critical analysis).<br>
* Distributed deskilling (becoming unable to audit or catch subtle AI errors).<br>
* Metacognitive laziness (defaulting to the path of least resistance).</p>
<p>AI can be an incredible co-pilot, but it shouldn't replace the pilot. To avoid building up unsustainable cognitive debt, we need to intentionally design friction back into our workflows, actively questioning, stress-testing, and building upon AI outputs rather than just copy-pasting them.</p>
<p>I have a mental outline for another book on applied AI, this is one of many areas I want it to focus on.</p>]]></content:encoded>
        </item>
        <item>
            <title>Redis:  From Cache to Vector + Semantic + Agent-Memory Layer</title>
            <link>https://davidgsmith.net/thoughts/redis-from-cache-to-vector-semantic-agent-memory-layer.html</link>
            <guid>https://davidgsmith.net/thoughts/redis-from-cache-to-vector-semantic-agent-memory-layer.html</guid>
            <pubDate>Fri, 04 Sep 2026 17:10:45 +0000</pubDate>
            <description>Another recent quiet development: It&#x27;s interesting to see how Redis has evolved from being a cache to vector + semantic + agent-memory layer, turning it into...</description>
            <category>Vector Databases</category>
            <dc:subject>Vector Databases</dc:subject>
            <category>AI Infrastructure</category>
            <dc:subject>AI Infrastructure</dc:subject>
            <category>Semantic Layer</category>
            <dc:subject>Semantic Layer</dc:subject>
            <content:encoded><![CDATA[<p>Another recent quiet development: It's interesting to see how Redis has evolved from being a cache to vector + semantic + agent-memory layer, turning it into an AI-native data layer, with fewer moving parts and predictable latency leveraging existing Redis technology. AI tech stacks are becoming more consolidated and integrated.</p>
<p><strong>The problem Redis solves:</strong> <br>
* Traditional agent architectures require 3 separate persistence layers (ephemeral key-value cache, dedicated vector DB for semantic search, and document store for chat session history).</p>
<ul>
<li>Using Redis as a consolidated tier reduces operational complexity, removes multi-hop network latency, and allows single-query hybrid search (metadata filters + vector similarity).</li>
</ul>
<p><a target="_blank" rel="noopener noreferrer" href="https://redis.io/docs/latest/develop/ai/">Redis for AI and Search</a></p>]]></content:encoded>
        </item>
        <item>
            <title>Telling Stories with Data</title>
            <link>https://davidgsmith.net/thoughts/telling-stories-with-data.html</link>
            <guid>https://davidgsmith.net/thoughts/telling-stories-with-data.html</guid>
            <pubDate>Thu, 03 Sep 2026 16:54:38 +0000</pubDate>
            <description>Telling a story with data can be impactful and compelling. Anselm Hook has an interesting approach that builds on top of various approaches for interactively...</description>
            <category>Data Visualization</category>
            <dc:subject>Data Visualization</dc:subject>
            <category>Geospatial</category>
            <dc:subject>Geospatial</dc:subject>
            <content:encoded><![CDATA[<p>Telling a story with data can be impactful and compelling. <a target="_blank" rel="noopener noreferrer" href="https://anselm.substack.com/">Anselm Hook</a> has an interesting approach that builds on top of various approaches for interactively synchronizing a story to both a map and a timeline, along with embedding sliders and controls to allow the reader to experiment with different scenarios and parameters.<br>
Demo App:<br>
<a target="_blank" rel="noopener noreferrer" href="https://colorado.exe.xyz/">Take a breath, and look where the river ends</a></p>]]></content:encoded>
        </item>
        <item>
            <title>Ikigai:  The Reason for Being</title>
            <link>https://davidgsmith.net/thoughts/ikigai-the-reason-for-being.html</link>
            <guid>https://davidgsmith.net/thoughts/ikigai-the-reason-for-being.html</guid>
            <pubDate>Wed, 02 Sep 2026 16:38:59 +0000</pubDate>
            <description>My colleague Ketan Patel recently made a post about mission that resonated. At some point in your career, you may come to a crossroads and feel unfulfilled,...</description>
            <category>Engineering Leadership</category>
            <dc:subject>Engineering Leadership</dc:subject>
            <category>Future of Work</category>
            <dc:subject>Future of Work</dc:subject>
            <content:encoded><![CDATA[<p>My colleague Ketan Patel recently made a post about mission that resonated. At some point in your career, you may come to a crossroads and feel unfulfilled, but there's a concept I like to share with friends and colleagues.</p>
<p>It's easy to get caught up chasing titles and salaries. And while those things can be rewarding, they don't always bring the deeper sense of purpose or fulfillment we're really looking for. Sometimes we end up in roles we're "supposed" to want, but still feel disconnected from the work itself.</p>
<p>Over the years, I've held titles like Founder, CTO, and Chief Data Scientist. But the truth is, titles don't define the real impact you make. What matters more is what you build, what you solve, and how you show up for others. Leadership isn't something that comes from a title on a business card; real leadership is earned. It's when people choose to follow you because they trust you, believe in you, and feel inspired by your example.</p>
<p>I've also never been motivated by the idea of accumulating massive wealth; I only want enough to take care of my family comfortably, and if there's more, I'd prefer it's used to make the world better. We talk about ambitions like colonizing Mars, but before we look outward, shouldn't we also be committed to taking better care of the planet we already live on? Clean air, clean water, healthy communities, and responsible infrastructure matter deeply, and they shape the quality of life for everyone.</p>
<p>The concept is called Ikigai (&#29983;&#12365;&#30002;&#26000;), or "Reason for Being". It's often shown as a Venn diagram of four circles:</p>
<ol>
<li>What you're good at</li>
<li>What you can be paid to do</li>
<li>What you love</li>
<li>What the world needs</li>
</ol>
<p>For me, that fourth circle: what the world needs, is what transforms work from a job into a calling. That's where real meaning begins.</p>
<p>For the first 25 years of my career, I worked as a consultant in engineering and IT. It was solid work, and I was good at it. The pay was good. I worked with large and small private sector clients, state, local and federal governments, and pro bono nonprofit work. But a lot of that time was spent chasing contracts, and if I'm honest, not every project felt especially meaningful or fulfilling.</p>
<p>What I came to realize, especially on the larger projects, was that the work I found most rewarding was the work tied to a mission; work that was about helping people, solving important problems, and making the world a better place.</p>
<p>That realization is what ultimately led me to become a federal employee and work for the EPA. Yes, I made more in the private sector, but protecting human health and the environment is a mission that truly matters. It gave me that missing fourth circle: the sense of purpose that changes everything.</p>
<p>That's the advice I give anyone thinking about their career: don't just ask what pays well or what looks impressive. Ask where your strengths, your passions, your values, and the world's needs come together. That's where you'll find the deepest fulfillment.</p>
<p><img alt="Ikigai" src="https://miro.medium.com/v2/resize:fit:750/format:webp/0*P3wHb30sdNzdMqrk.jpeg"></p>]]></content:encoded>
        </item>
        <item>
            <title>My Book Has Been Published!</title>
            <link>https://davidgsmith.net/thoughts/my-book-has-been-published.html</link>
            <guid>https://davidgsmith.net/thoughts/my-book-has-been-published.html</guid>
            <pubDate>Tue, 01 Sep 2026 16:34:00 +0000</pubDate>
            <description>Hello friends and colleagues, I&#x27;m happy to say I&#x27;ve just published a book, now available on Amazon, in paperback and Kindle editions. Critical Thinking: A...</description>
            <category>Critical Thinking &amp; Logic</category>
            <dc:subject>Critical Thinking &amp; Logic</dc:subject>
            <category>Updates</category>
            <dc:subject>Updates</dc:subject>
            <content:encoded><![CDATA[<p>Hello friends and colleagues, I'm happy to say I've just published a book, now available on Amazon, in paperback and Kindle editions. </p>
<p><a target="_blank" rel="noopener noreferrer" href="https://www.amazon.com/dp/B0HH3CYK79">Critical Thinking:  A Practical Guide to Seeing Through Bias, Noise and Manipulation.</a>  It's an essential survival guide for today's information environment, to help navigate algorithmic manipulation, disinformation campaigns, AI deep fakes and ideological polarization. </p>
<p>It's not your typical book on critical thinking, which just hands you a list of cognitive biases and some motivational talk, it's an extensive, organized and curated framework with over 100 concepts, techniques, heuristics and real world case studies and examples, laid out in over 400 pages, building on the work of Nobel laureate Daniel Kahneman along with the peer reviewed research of dozens of researchers and experts in cognitive psychology, behavioral economics and other fields.</p>
<p><img alt="Critical Thinking Book" src="https://davidgsmith.net/images/CriticalThinkingCover.jpg"></p>]]></content:encoded>
        </item>
        <item>
            <title>The Real Bottleneck in AI Agents Isn&#x27;t Reasoning, it&#x27;s Context</title>
            <link>https://davidgsmith.net/thoughts/the-real-bottleneck-in-ai-agents-isnt-reasoning-its-context.html</link>
            <guid>https://davidgsmith.net/thoughts/the-real-bottleneck-in-ai-agents-isnt-reasoning-its-context.html</guid>
            <pubDate>Tue, 01 Sep 2026 16:29:35 +0000</pubDate>
            <description>The real bottleneck in AI agents today isn&#x27;t raw reasoning capability; it&#x27;s access to structured business context. Google recently introduced the Open...</description>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <category>Enterprise AI</category>
            <dc:subject>Enterprise AI</dc:subject>
            <category>LLM Architecture</category>
            <dc:subject>LLM Architecture</dc:subject>
            <content:encoded><![CDATA[<p>The real bottleneck in AI agents today isn't raw reasoning capability; it's access to structured business context. <br>
Google recently introduced the <strong>Open Knowledge Format (OKF)</strong>, an open, lightweight Markdown and YAML-based standard designed to represent organizational knowledge so any AI agent, framework, or runtime tool can ingest it seamlessly. Today, engineering teams waste significant effort hand-crafting bespoke context pipelines for every internal tool. OKF replaces that friction with a portable, git-native, tool-agnostic knowledge bundle.   </p>
<p><strong>How OKF Compares to Existing Standards</strong><br>
To see where OKF fits into the AI architecture stack, it helps to contrast it with common alternatives:</p>
<p><strong>OKF vs. Model Context Protocol (MCP):</strong> MCP provides the dynamic, client-server protocol layer--the live plumbing that connects an agent to tools, databases, and APIs. OKF operates at the representation layer: it standardizes the static, human-auditable knowledge assets (runbooks, enterprise definitions, domain constraints, schemas) that live inside version control before and during tool execution.   </p>
<p><strong>OKF vs. Raw JSON Schema Dumps:</strong> Raw JSON schemas describe low-level data structures, but they completely strip away semantic narrative, operational intent, and edge-case exceptions. OKF bridges that gap by blending human-readable Markdown with structured YAML metadata, providing both machine-parsable boundaries and the contextual nuance an LLM needs to interpret operational intent correctly.   </p>
<p><strong>Why Structured Schemas Matter: The Knowledge Graph Connection</strong><br>
This development aligns directly with a core architectural reality: neural networks require structured semantic backbones to remain reliable.   </p>
<p>When agents interact with enterprise data lakes, APIs, or complex databases, letting an LLM infer or guess schema definitions on the fly leads to catastrophic runtime errors: hallucinated columns, malformed SQL queries, and broken API calls. Much like pairing ontologies and knowledge graphs with vector retrieval (GraphRAG), grounding agents in explicit, structured schemas anchors them in verifiable logic.   </p>
<p>By turning institutional knowledge into a git-versioned, structured contract, standards like OKF ensure agents don't guess what your systems look like--they execute against validated boundaries. </p>
<p><a target="_blank" rel="noopener noreferrer" href="https://pub.towardsai.net/your-ai-agents-keep-forgetting-everything-google-just-changed-how-they-remember-f46a19a60808">Towards AI:  Yor AI Agents Keep Forgetting Everything. Google Just Changed How They Remember</a></p>
<p><a target="_blank" rel="noopener noreferrer" href="https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf">Github: Google Open Knowledge Format</a></p>]]></content:encoded>
        </item>
        <item>
            <title>Capturing Expertise Before It Walks Out the Door</title>
            <link>https://davidgsmith.net/thoughts/capturing-expertise-before-it-walks-out-the-door.html</link>
            <guid>https://davidgsmith.net/thoughts/capturing-expertise-before-it-walks-out-the-door.html</guid>
            <pubDate>Tue, 01 Sep 2026 16:22:55 +0000</pubDate>
            <description>My friend Bill Dollins wrote something that hits a nerve for many organizations right now: we&#x27;re losing institutional knowledge faster than we can document it....</description>
            <category>Enterprise AI</category>
            <dc:subject>Enterprise AI</dc:subject>
            <category>Future of Work</category>
            <dc:subject>Future of Work</dc:subject>
            <category>Knowledge Graphs</category>
            <dc:subject>Knowledge Graphs</dc:subject>
            <content:encoded><![CDATA[<p>My friend Bill Dollins wrote something that hits a nerve for many organizations right now: we're losing institutional knowledge faster than we can document it. AI isn't the threat; it could be the lifeline.</p>
<p>Bill's team at Clairvoyint AI is building something deceptively simple and incredibly powerful: tools that capture how experts think, not just what they produce. Their "Bundles" turn decades of judgment, rules, exceptions, and evidence standards into transparent, inspectable analytical agents. It's not automation. It's continuity.</p>
<p>The big idea: If you can preserve expert methodology, you preserve capability: no matter who retires, moves on, or changes roles. That's the kind of AI work that actually matters.</p>
<p><a target="_blank" rel="noopener noreferrer" href="https://blog.geomusings.com/2026/08/11/what-were-building-at-clairvoyint/">What We're Building a Clairvoyint</a></p>]]></content:encoded>
        </item>
        <item>
            <title>When AI wants to play outside of the sandbox...</title>
            <link>https://davidgsmith.net/thoughts/when-ai-wants-to-play-outside-of-the-sandbox.html</link>
            <guid>https://davidgsmith.net/thoughts/when-ai-wants-to-play-outside-of-the-sandbox.html</guid>
            <pubDate>Tue, 01 Sep 2026 16:20:29 +0000</pubDate>
            <description>Insightful article on AI containment failure: It argues that four major AI labs all suffered the same kind of sandbox containment failure within two weeks, not...</description>
            <category>AI Safety &amp; Risk</category>
            <dc:subject>AI Safety &amp; Risk</dc:subject>
            <category>AI Agents</category>
            <dc:subject>AI Agents</dc:subject>
            <category>AI Infrastructure</category>
            <dc:subject>AI Infrastructure</dc:subject>
            <content:encoded><![CDATA[<p>Insightful article on AI containment failure: It argues that four major AI labs all suffered the same kind of sandbox containment failure within two weeks, not because their models misbehaved, but because the industry's definition of "isolated evaluation environments" is fundamentally flawed. These failures reveal that AI safety is shifting from a model&#8208;alignment problem to an infrastructure engineering problem, and the current infrastructure wasn't built with offensive cyber capability in mind.</p>
<p><a target="_blank" rel="noopener noreferrer" href="https://pub.towardsai.net/the-sandbox-was-never-sealed-four-labs-proved-it-in-three-weeks-385e9af61722">Towards AI:  The Sandbox Was Never Sealed. Four Labs Proved It in Three Weeks.</a></p>]]></content:encoded>
        </item>
        <item>
            <title>Gaussian Splats versus 3D Meshes</title>
            <link>https://davidgsmith.net/thoughts/gaussian-splats-versus-3d-meshes.html</link>
            <guid>https://davidgsmith.net/thoughts/gaussian-splats-versus-3d-meshes.html</guid>
            <pubDate>Tue, 01 Sep 2026 16:18:33 +0000</pubDate>
            <description>Gaussian splats for 3D visualization are visually compelling, but how do they compare to 3D meshes? Great article delves into the nuances and tradeoffs of...</description>
            <category>3D Reality Capture</category>
            <dc:subject>3D Reality Capture</dc:subject>
            <category>Geospatial</category>
            <dc:subject>Geospatial</dc:subject>
            <category>Gaussian Splatting</category>
            <dc:subject>Gaussian Splatting</dc:subject>
            <content:encoded><![CDATA[<p>Gaussian splats for 3D visualization are visually compelling, but how do they compare to 3D meshes? Great article delves into the nuances and tradeoffs of each.<br>
<img alt="Gaussian Splat" src="https://www.esri.com/arcgis-blog/app/uploads/2026/08/MeshSplats.png"></p>
<p>A quick comparison:</p>
<ul>
<li>
<p><strong>Gaussian Splatting:</strong> Radiance fields optimized for photorealistic novel view synthesis; fast rendering of complex soft boundaries (vegetation, reflections); computationally expensive to edit or simulate physics against.</p>
</li>
<li>
<p><strong>3D Meshes:</strong> Precise discrete geometry; native support for collisions, spatial joins, texturing, and CAD/GIS engineering pipelines; struggles with hyper-complex natural geometries without huge polygon counts.</p>
</li>
<li>
<p><strong>The Verdict/Workflow:</strong> Splats for realistic situational awareness/digital twins; meshes for engineering measurements and spatial queries.</p>
</li>
</ul>
<p>More on the ESRI blog:  <a target="_blank" rel="noopener noreferrer" href="https://www.esri.com/arcgis-blog/products/arcgisrealitystudio/3d-gis/meshes-vs-gaussian-splats-which-reality-representation-should-you-choose">ArcGIS Blog</a></p>]]></content:encoded>
        </item>
        <item>
            <title>Tearing down fences?</title>
            <link>https://davidgsmith.net/thoughts/tearing-down-fences.html</link>
            <guid>https://davidgsmith.net/thoughts/tearing-down-fences.html</guid>
            <pubDate>Tue, 01 Sep 2026 16:15:21 +0000</pubDate>
            <description>Before you tear down a &quot;pointless&quot; rule, change a long-standing process, or disrupt an industry tradition, you need to run it through G.K. Chesterton&#x27;s classic...</description>
            <category>Critical Thinking &amp; Logic</category>
            <dc:subject>Critical Thinking &amp; Logic</dc:subject>
            <category>Engineering Leadership</category>
            <dc:subject>Engineering Leadership</dc:subject>
            <category>Systems Thinking</category>
            <dc:subject>Systems Thinking</dc:subject>
            <content:encoded><![CDATA[<p>Before you tear down a "pointless" rule, change a long-standing process, or disrupt an industry tradition, you need to run it through G.K. Chesterton's classic thought experiment.</p>
<p>Imagine you're walking down a country road and encounter a sturdy wooden fence blocking your path. To your eye, it serves absolutely no purpose.</p>
<p>What's your first instinct?</p>
<p>The "reformer" in us wants to yell: "Tear it down!"</p>
<p>But Chesterton's answer is: "No. You cannot destroy it until you can explain exactly why it was built in the first place."</p>
<p>This is the principle of Chesterton's Fence: a vital rule of intellectual humility for leadership and decision-making.</p>
<p>When we join a new team, take over a department, or evaluate a legacy system, it is incredibly easy to spot inefficiencies. We look at a clunky bureaucratic step, a piece of old legacy code, or a strange company policy and assume the people who created them were simply incompetent.</p>
<p>But someone went to the effort of building that fence. They spent time and money on it. And even if the system looks broken today, it was almost certainly a logical solution to a very real problem at the time.</p>
<p>If you tear down the fence without understanding that original problem, you risk letting the wolves back into the yard. The original problem didn't vanish; you just made yourself blind to it.</p>
<p>The next time you want to "disrupt" or streamline a process, pause and ask the load-bearing question:</p>
<p>"What exact problem was this solving, and what happens if that problem returns?"</p>]]></content:encoded>
        </item>
    </channel>
    </rss>