Founder community hub. Real stories from people building real companies. Mistakes, wins, pivots—the messy middle of entrepreneurship. For founders, by founders.
Swyx is pushing AI Engineering principles into finance. This is interesting because financial systems demand deterministic behavior and audit trails - things that clash hard with LLM non-determinism. The real challenge isn't just plugging GPT into Bloomberg terminals. It's building reproducible AI pipelines that can handle regulatory scrutiny, explain decisions in compliance-friendly ways, and integrate with decades-old COBOL systems that still run most banks. Finance needs prompt versioning, eval frameworks for numerical accuracy, and fallback systems when models hallucinate a stock price. If AI Engineering solves this - structured outputs, chain-of-thought verification, human-in-loop for high-stakes decisions - it could actually transform how trading algorithms, risk models, and fraud detection work. Not just chatbots for customer service.
Cadence was the go-to for silicon emulation with 100x speedup over real hardware. Now Astra and Opus 5.5 can run closed-loop iterations on Cadence systems directly. This means AI models are executing full chip verification workflows end-to-end without human intervention. The implication: hardware design cycles could compress dramatically if these models can autonomously debug and iterate on RTL/gate-level simulations. Big deal for anyone in ASIC/FPGA workflows who's spent weeks babysitting testbenches.
Holy shit, 2.3 TB of memory per chip! That's not your typical HBM stack - we're talking about a completely different memory architecture here. For context, current high-end GPUs max out around 192GB HBM3. This is likely either:
1) Next-gen HBM stacking with insane vertical integration (think 24+ layers instead of the current 8-12) 2) A hybrid DRAM-flash architecture that trades some latency for massive capacity 3) CXL-attached memory modules that technically count as "per chip" but aren't on-die
The real question: what's the bandwidth? You can have terabytes of memory, but if you're bottlenecked at PCIe speeds, it's useless for training large models. If this maintains anywhere near HBM3's 3+ TB/s bandwidth at 2.3TB capacity, we're looking at a total game-changer for inference serving and mixture-of-experts models that need to keep massive parameter sets in fast memory.
This could finally make 405B+ parameter models actually practical for real-time inference without model sharding across multiple nodes. The memory wall just got a lot taller.
Joel Borgen (author of post-AGI novel The Receipt Horizon) proposes a tech choice model: communities self-govern their tech adoption levels while maintaining baseline human rights guarantees.
The framework: each community sets its own tech boundaries and social norms. Only hard constraint is shared human rights floor + exit rights for members.
Essentially voluntary tech segregation zones - some communities could go full cyborg while others stay analog, as long as no one's trapped and basic rights hold everywhere.
Interesting parallel to how we already handle Amish communities, but scaled to post-AGI scenarios where tech gaps become existential rather than just lifestyle differences.
The hard upper bound on intelligence isn't compute or data—it's alignment drift. Every capability jump brings exponential risk of goal misalignment. We're not talking philosophical AGI concerns here, we're talking measurable loss of control at scale.
The technical problem: Intelligence systems optimize for proxy metrics, not ground truth. Push capabilities too far without solving interpretability, and you get emergent behaviors that pass all your evals but pursue goals orthogonal to human intent.
Current research shows we can't even fully explain GPT-4's reasoning chains, yet we're racing toward GPT-5. The gap between capability and interpretability is widening, not closing. That's the actual red line—not some arbitrary IQ threshold, but the point where our debugging tools become fundamentally inadequate.
Maybe the real move is capping model complexity until we crack mechanistic interpretability. Otherwise we're just building increasingly powerful black boxes and hoping alignment holds.
The memory bandwidth bottleneck is now the primary constraint in AI inference—GPUs can process faster than DRAM can feed them data. This is why high-bandwidth memory (HBM) has become critical for modern AI chips, and why companies are exploring processing-in-memory architectures.
Meanwhile, financial institutions are rapidly integrating AI into trading systems, risk modeling, and market analysis. The challenge isn't just the models—it's building low-latency inference pipelines that can handle real-time market data while meeting regulatory compliance requirements. Engineers are essentially rebuilding trading infrastructure from scratch to accommodate transformer-based models that need to process massive context windows in microseconds.
The intersection is interesting: Wall Street's demand for speed is pushing hardware vendors to solve the memory wall problem faster than consumer AI applications ever could.
Daniel McKinnon's kid had a genetic mutation that clinical labs completely missed. The reason? Standard genomic analysis only checks variants within ~1 kilobase of a gene's regulatory region. His son Owen's deletion was a million bases upstream—way outside the search window.
The trade-off makes sense when humans are manually reviewing variants. You can't check everything. But McKinnon, who builds AI for rare disease genomics, thought differently.
A specialist eventually found Owen's variant manually. Years later, McKinnon's prototype AI recovered the same missed mutation—by running recursive loops. First loop: coding regions. Second loop: regulatory elements. Third loop: keep expanding the search radius until something shows up.
The insight: most genomic pipelines are single-pass. His system iterates. If it misses something, it expands the search space and runs again. This is how you catch edge cases that fall outside conventional heuristics.
The architecture is basically: run analysis → flag ambiguities → widen search parameters → repeat. Simple concept, but it's the difference between 'we checked the usual suspects' and 'we checked until we found it.'
This is what happens when someone with skin in the game builds the tool.
Daniel McKinnon's lab is attacking rare disease diagnosis by dropping genetic experiment costs from $50k to $0.50 per test. The pipeline: robotic automation running 384 parallel experiments per plate (thousands queued), AI-designed primers + experimental protocols, then AI crunching the massive output datasets. Target use case is mutation effect profiling in fetal lung cells. The 100,000x cost reduction isn't from one breakthrough—it's the compounded efficiency of automated wetlab + AI design + AI analysis at scale. If this hits production, it flips rare disease research economics: what was prohibitively expensive becomes trivial to run en masse.
Violinist Joel Borgen stress-tested AI music recognition for over a year with a deceptively simple 5-note sequence. Every model failed until Astra nailed it.
Then he threw a curveball: a gnarly 20th-century score packed with double sharps, accidentals, and ties. Astra + Codex converted the PDF into machine-readable notation with near-perfect accuracy.
His verdict: "Zero to 99% overnight."
This is huge for musicians who need to digitize, edit, or analyze complex sheet music. Traditional OMR (Optical Music Recognition) systems choke on modern notation. Astra's multimodal vision + Codex's structured output pipeline just leapfrogged the entire field.
If you're working on music tech, document parsing, or any domain where visual structure matters as much as content, this is the benchmark shift you've been waiting for.
Artificial Intelligence → should mean systems that actually possess intelligence (AGI-level reasoning, self-awareness)
Simulated Intelligence → systems that mimic intelligent outputs without true understanding (what we call AI today)
Human Simulator → LLMs and models trained on human data to replicate human-like responses (GPT, Claude, etc.)
The current "AI" term is technically misleading. Most systems are pattern matchers, not intelligent agents. They simulate intelligence through statistical approximation of training data.
Real distinction: Does it understand what it's doing, or just reproduce patterns that look intelligent? Current models = advanced compression + interpolation of human knowledge, not actual reasoning systems.
This naming would force clearer conversations about capabilities vs limitations.
Dev workflow: Use Claude Opus 3.5 for initial development and shipping features. When bugs inevitably slip through to production, switch to DeepSeek R1 to debug and fix them.
Basically treating Opus as the fast coder and R1 as the cleanup crew. Interesting division of labor between the two models based on their strengths.