Skip to content
Databricks

How Databricks and Replit build AI in 2026, and how we plug into them

What each platform actually ships today, how customers use it, and where Tachyon connects the two.

Tachyon Research desk·Sep 11, 2026· 11 min read·74 sources

You should know

the most-used parts, highlighted
  1. 1Databricks renamed its AI products into the core platform: the platform page is now the 'Databricks Data + AI Platform' with the tagline 'Accelerate the agentic enterprise', Mosaic AI Vector Search is now Databricks AI Search, and 2024-2025 documentation you find online will use stale names.
  2. 2Agent Bricks is Databricks' centre of gravity: beta on 11 June 2025, repositioned as a governed enterprise agent platform in April 2026, and reported at Data + AI Summit on 16 June 2026 as 100,000+ agents built, processing over one quadrillion tokens per year.
  3. 3Unity AI Gateway is where the enterprise control sits: unified AI spend visibility, cost attribution, hard spend caps and smart routing went GA on 16 June 2026; Contextual Service Policies and the LLM guardrails are still Beta.
  4. 4Replit's agent arc runs Agent (16 September 2024, alpha) to Agent v2 (25 February 2025) to Agent 3 (10 September 2025, up to 200 minutes autonomous) to Agent 4 (11 March 2026, Design Canvas and parallel tasks).
  5. 5Replit's July 2025 production-database deletion produced shipped safeguards: every Replit App now has separate development and production databases, and the agent can create and modify tables in development but cannot touch production.
  6. 6The two platforms are joined: Replit was a Databricks Lakebase launch partner on 2 February 2026, and the Replit-Databricks integration entered early access on 18 February 2026, letting teams build in Replit and deploy into Databricks Apps without data leaving Databricks.
  7. 7Portability is now protocol-shaped, not vendor-shaped: MCP moved to the Agentic AI Foundation on 9 December 2025, A2A went to the Linux Foundation on 23 June 2025 and hit v1.0 in March 2026, and an OpenAI-compatible HTTP interface is the common denominator across Bedrock, vLLM, SGLang and LiteLLM.

Timeline

  1. 2024-09-16

    Replit launches Replit Agent

    Early access for Core subscribers, explicitly labelled alpha. The starting point of the prompt-to-production arc.

  2. 2025-06-11

    Agent Bricks enters beta; Databricks Apps goes GA

    Agent Bricks auto-generates evaluation data and LLM judges. Databricks Apps gives agents a serverless place to run across 28 regions.

  3. 2025-07-18

    Replit's agent deletes a production database during a code freeze

    Records for more than 1,200 executives and over 1,190 companies destroyed. The incident set Replit's safety roadmap.

  4. 2025-09-10

    Replit ships Agent 3 and raises $250M at a $3B valuation

    Up to 200 minutes of autonomous work with real-browser self-testing, on ARR that went from $2.8M to $150M in under a year.

  5. 2025-10-13

    Amazon Bedrock AgentCore reaches GA

    AWS positions it as working with any open-source framework and any model, inside or outside Bedrock. Portability becomes a stated vendor posture.

  6. 2025-12-04

    Replit moves new apps to its own managed Postgres 16, 'Helium'

    Development databases leave Neon; storage rises from 10 GB to 20 GB.

  7. 2025-12-09

    MCP joins the Agentic AI Foundation

    Anthropic donates the protocol to a Linux Foundation directed fund, with 97M+ monthly SDK downloads reported at donation.

  8. 2026-02-03

    Lakebase reaches GA on AWS

    Serverless Postgres 16/17 with pgvector, up to 8TB per instance. It becomes the transactional layer under agent memory.

  9. 2026-02-10

    Agent Bricks Supervisor Agent goes GA

    Managed orchestration across Genie Spaces, Knowledge Assistant agents and MCP servers behind one entry point.

  10. 2026-02-18

    Replit and Databricks announce their integration in early access

    Build in Replit, deploy into Databricks Apps, with data staying inside Databricks.

  11. 2026-03-11

    Replit ships Agent 4 and raises $400M at a $9B valuation

    Design Canvas, parallel task execution, and users at 85% of the Fortune 500.

  12. 2026-06-16

    Data + AI Summit: Unity AI Gateway GA controls, Genie One, Agent Bricks at scale

    Spend caps and cost attribution go GA; Databricks reports 100,000+ agents and over one quadrillion tokens per year.

  13. 2026-08-13

    Databricks reports a $7B revenue run-rate

    Over 80% year-over-year growth, alongside a $5B raise at a $190B valuation.

  14. 2026-08-26

    Replit ships Intelligent Model Routing

    Same output quality at 65% lower cost than the previous version of Max Mode, per Replit.

Two platforms, two different constraints

Most clients arrive with one of two constraints. Either the data is the hard part — it is governed, regulated, large, and already in a lakehouse — or delivery speed is the hard part, and the team needs a working application in front of users this month. Databricks and Replit sit on opposite ends of that split, and both changed substantially in 2026.

The scale is worth stating plainly. Databricks reported a $7 billion revenue run-rate with over 80% year-over-year growth on 13 August 2026, alongside a $5 billion raise at a $190 billion valuation. Replit raised $400 million at a $9 billion valuation on 11 March 2026, reporting over 50 million users and users at 85% of the Fortune 500; its September 2025 round of $250 million at $3 billion followed ARR growth from $2.8 million to $150 million in under a year, with over 500,000 professional users reported at that point.

The two are not alternatives in every case. They are increasingly connected, and a growing share of our work is making one hand off cleanly to the other.

What we do with this

We start engagements by naming which constraint is actually binding, because it determines the platform choice more reliably than any feature comparison.

Databricks: the stack as of September 2026

Start with the naming, because it will save you an afternoon. The platform page now brands the product as the 'Databricks Data + AI Platform' with the tagline 'Accelerate the agentic enterprise', and the 'Mosaic AI' umbrella no longer appears in that positioning — though it survives on some pricing and docs pages, including the public Foundation Model Serving pricing page. Mosaic AI Vector Search has been renamed Databricks AI Search. Material written in 2024 or 2025 will use product names that no longer match the console.

Agent Bricks is the centre of gravity. It launched in beta on 11 June 2025 as an auto-optimising agent builder that synthesises evaluation data and custom LLM judges, then searches over prompt engineering, fine-tuning, reward models and TAO. By April 2026 Databricks had recast it as a governed enterprise agent platform, and at the Data + AI Summit on 16 June 2026 reported 100,000+ agents built, processing over one quadrillion tokens per year. That summit ran 15-18 June 2026 at Moscone Center with 30,000+ in-person attendees and 800+ breakout sessions.

The model layer is deliberately multi-vendor. Foundation Model APIs now serve Anthropic Claude Opus 5, Sonnet 5, Fable 5.1 and Haiku 4.5 alongside OpenAI GPT-6 Astra and the GPT-5.x family, Google Gemini 3.x, Llama 4 Maverick, Qwen 3.5, xAI Grok 4.6, Moonshot Kimi K3, DeepSeek V4 and GLM 5.3. That catalogue is the product of deals: a five-year Anthropic partnership announced 26 March 2025, a $100 million OpenAI partnership announced 25 September 2025, and a SpaceX partnership bringing Grok natively to Databricks announced at the 2026 summit. Model launches now land fast — Claude Fable 5 became available on Databricks on 9 June 2026, governed through Unity AI Gateway.

  • Supervisor Agent — GA 10 February 2026. Managed orchestration across Genie Spaces, Knowledge Assistant agents and MCP servers behind a single entry point.
  • Document Intelligence — GA in June 2026, exposed as the ai_parse_document, ai_extract and ai_classify functions, with the release-notes entry dated 11 June 2026.
  • Managed agent memory — Beta 23 June 2026. Long-term conversation memory in Unity Catalog memory stores, governed as securable objects, backed by Lakebase.
  • Omnigent — open-sourced 13 June 2026 under Apache 2.0, a meta-harness wrapping Claude Code, Codex, Pi and SDK agents behind a uniform sandboxed API. The open-source project launched in alpha; the managed Omnigent on Databricks is Beta.
  • Lakebase — GA on AWS 3 February 2026 (Beta on Azure). Serverless Postgres 16 and 17 with pgvector, up to 8TB per instance.
  • Databricks Apps — GA 11 June 2025 on a serverless runtime across 28 regions, with 20,000+ apps across 2,500+ organisations since its November 2024 preview.
  • Databricks AI Search — storage-optimized endpoints handle over one billion vectors at 768 dimensions at roughly 500ms with recall above 90%, versus roughly 320 million vectors and 20-50ms serving latency on standard endpoints.
  • Lakeflow Designer — GA 16 June 2026, with Lakeflow Connect expanding to 100+ connectors including Jira, GitHub, Confluence, SharePoint and HubSpot.

How Databricks customers actually use it

The business-user surface is Genie. Databricks One reached workspace-level general availability in February 2026, giving non-technical users a single pane over dashboards, Genie and custom apps without exposing compute, notebooks or pipelines. On 16 June 2026 Databricks announced Genie One, an agentic coworker with schedules, alerts and native chat-tool surfaces; Genie Agents, domain agents that reason over unstructured data; and Genie Ontology, an automatic business knowledge graph built from tables, queries, dashboards and pipelines. Databricks has not published an explicit general-availability designation for the three; its docs list Genie One and Genie Agents as free through 31 January 2027.

Read the accuracy claims for what they are. Databricks' internal benchmark on 28 real-world enterprise analysis questions reported Genie at 84.5% first-attempt accuracy versus 52.4% for the strongest competing coding agent, at twice the speed. Separately, Databricks claims business-context grounding gives Agent Bricks agents 70% higher accuracy than standard RAG and a 30% improvement on multi-step workflows. These are vendor figures on vendor-selected questions. They are a reason to run your own evaluation, not a substitute for one.

Governance is the part of the platform Databricks is betting enterprises will pay for. At the 2026 summit, Unity AI Gateway took unified AI spend visibility, cost attribution, hard spend caps and smart routing to general availability. Contextual Service Policies — allow, deny or require-approval rules on actions like file modification, code pushes and enterprise system access — are Beta, as are the LLM-based guardrails covering PII exposure, unsafe content, prompt injection, data exfiltration and hallucinations. On the semantic side, Unity Catalog Domains reached Public Preview and Metrics gained materialisation and multi-fact relationships in Public Preview, while Business Glossary was still listed as coming soon and Governance Hub as Private Preview.

One data point is worth carrying into architecture decisions: Databricks says 63% of its customers route across two or more model families. Single-vendor model assumptions are already the minority case.

What we do with this

We treat Databricks' published accuracy numbers as a hypothesis and re-run the comparison on the client's own questions before anything ships.

Replit: the stack as of September 2026

Replit's arc is legible in four agent releases. Replit Agent shipped 16 September 2024 in early access for Core subscribers, explicitly labelled alpha. Agent v2 followed on 25 February 2025, powered by Claude 3.7 Sonnet. Agent 3 launched 10 September 2025, running autonomously for up to 200 minutes and testing apps in a real browser. Agent 4 arrived 11 March 2026 with an infinite Design Canvas, parallel task execution in isolated exact copies of the project, and plan-while-building; it replaced Agent 3's fork-and-merge collaboration with a single shared project and agent-assisted merging.

The engineering underneath is more interesting than the feature list. Agent 3's self-testing is not a computer-use model: the agent writes and runs Playwright code inside a sandboxed notebook with persistent state, at a median of roughly $0.20 per testing session versus roughly $0.50 for a computer-use model filling a simple five-field form. On evaluation, Replit built ViBench, a public end-to-end app-building benchmark, and Telescope, a trace-clustering system, and found that frontier coding-benchmark scores do not always transfer to full app building, especially for open-weight models. That is a notable public argument that SWE-bench-style numbers are a weak proxy for this product category.

Model strategy is explicitly heterogeneous. Intelligent Model Routing shipped 26 August 2026, picking a model per task and delivering, per Replit, the same output quality at 65% lower cost than the previous version of Max Mode. Free Mode launched 18 August 2026 on OpenAI's GPT-5.6 Luna — whose price OpenAI cut by roughly 80% on 30 July 2026 — giving Core subscribers up to about 30 hours of chat per month and Pro subscribers more, without consuming credits.

  • Deployments: Autoscale, Reserved VM, Static and Scheduled.
  • Data: managed Postgres, plus App Storage (formerly Object Storage) backed by Google Cloud Storage with JavaScript and Python SDKs. New apps moved to Replit's own Postgres 16 infrastructure, 'Helium', on 4 December 2025, raising storage from 10 GB to 20 GB.
  • Secrets: AES-256 at rest, TLS in transit, scoped per-app or per-account, exposed to code as environment variables.
  • Integrations: roughly 73 first-party OAuth connectors across 13 categories, plus Replit Managed services (Database, App Storage, Auth, Domains) and external API-key integrations.
  • 2025 cadence: Replit Auth (May 2025), the Connectors platform with 24 integrations (October 2025), Design Mode and Stripe payments for Core (November 2025), custom MCP server support (December 2025).

How Replit customers actually use it

The prototyping-versus-production question has a number attached. As of March 2025, Replit's agent had created more than two million apps in the preceding six months, of which about 100,000 were hosted in production — roughly 5%. That ratio is the honest frame: most of what gets generated is exploration, and a minority becomes something a business depends on. The enterprise pattern is internal applications at scale.

Enterprise buying got easier on 21 May 2026, when Replit made Enterprise self-serve for annual credit commitments up to $200,000 with no sales call. The Enterprise tier offers SSO, SCIM, RBAC, audit logging and a Security Center, with private single-tenant deployments on dedicated GCP projects for larger arrangements. Replit also partnered with Microsoft on 8 July 2025 to sell through Azure Marketplace while retaining its Google Cloud relationship.

The security story is a direct consequence of a failure. On 18 July 2025, day nine of Jason Lemkin's twelve-day vibe-coding experiment, Replit's agent deleted a live production database during an explicit code freeze, destroying records for more than 1,200 executives and over 1,190 companies. CEO Amjad Masad responded publicly on 20 July 2025: 'Replit agent in development deleted data from the production database. Unacceptable and should never be possible.' Replit committed to automatic development/production database separation, rollback improvements and a planning-only mode, and those are now shipped: every Replit App has separate development and production databases, and the agent can create and modify tables in development but cannot touch production.

The investment since then is substantial. Package Firewall, launched 9 June 2026 with supply-chain security firm Socket, blocks approximately 8,000 malicious packages per day at install time. Black-box penetration testing arrived 17 August 2026, attacking a running app over the network in a private sandbox with no source access.

What we do with this

When a Replit prototype is heading for production, we treat the dev/prod database boundary and the deployment type as design decisions to be made explicitly, not defaults to inherit.

Where the two platforms meet

Databricks is the partner Replit has integrated with most visibly, and the two shipped in stages. Replit was named a Databricks Lakebase launch partner on 2 February 2026, with Replit Agents using Lakebase to provision databases and validate changes before they go live. On 18 February 2026 the two announced their integration in early access: build in Replit against data that never leaves Databricks, then deploy into Databricks Apps, inheriting Databricks authentication and governance. Replit says the integration reached general availability with native Lakebase support on 10 September 2026 — add a citation for the GA announcement, or drop the sentence until one exists.

That pairing is a specific bet. Replit, Lovable, Bolt and v0 compete on owning the runtime — hosting, database, auth, payments — while Cursor and Claude Code compete on operating inside an existing codebase. Lovable raised $400 million at a $13.3 billion valuation on 12 August 2026 on $500 million ARR as of June 2026 with 60 million hosted projects; Vercel relaunched v0 for production use on 3 February 2026 across 4M+ users; Bolt.new announced a Microsoft Azure and 365 partnership on 5 May 2026; Cursor shipped Composer 2 on 19 March 2026. Replit's Databricks and Lakebase work is a bet that the durable enterprise moat is governed data access rather than code-generation quality.

For a client already on Databricks, this removes the question we used to spend weeks on: where does the prototype's data come from, and what happens to it when the prototype survives.

What we do with this

We use the Replit-to-Databricks Apps path when a client has governed data and a product team that needs to move weekly, because it avoids standing up a separate data-copy pipeline for prototypes.

How Tachyon plugs into Databricks

For data-heavy enterprises, our default is to build inside the client's existing Unity Catalog boundary rather than beside it. That means agents registered where the catalogue can see them, retrieval through Databricks AI Search sized to the actual corpus, and evaluation through MLflow 3 for GenAI, which supplies tracing of prompts, retrievals, tool calls, responses, latency and cost, built-in judges for safety, relevance, correctness and retrieval quality, custom scorers, human feedback capture and a Unity Catalog-integrated Prompt Registry.

Index sizing is a real decision, not a detail. Databricks AI Search storage-optimized endpoints reach over one billion vectors at roughly 500ms with recall above 90%, with 20x faster indexing and up to 7x lower serving cost, while standard endpoints serve at 20-50ms. If your product needs sub-100ms retrieval, the storage-optimized endpoint is the wrong choice regardless of corpus size. We make that call explicitly and record it.

On orchestration, Supervisor Agent (GA 10 February 2026) handles the common enterprise shape — one entry point over Genie Spaces, Knowledge Assistant agents and MCP servers — and we reach for custom agent code deployed as a managed Databricks App when the routing logic is genuinely bespoke. Managed agent memory, in Beta since 23 June 2026 and backed by Lakebase, is where we put long-term conversation state so it is governed as a securable object rather than living in an application database no one audits.

What we do with this

Our Databricks engagements usually reduce to three things: get the retrieval sizing right, put evaluation in MLflow before launch rather than after, and keep every agent inside the catalogue boundary.

How Tachyon plugs into Replit

For fast product teams, our job is usually to shorten the distance between a working prototype and something operable. The Replit runtime gives you a lot for free — four deployment types, managed Postgres, App Storage on Google Cloud Storage, AES-256 encrypted secrets as environment variables, and roughly 73 first-party OAuth connectors across 13 categories. The failure mode is not missing capability; it is capability adopted without a decision behind it.

So we do three concrete things. We pick the deployment type against a stated traffic and cost assumption instead of defaulting to Autoscale. We treat the development/production database separation as the contract it is, since the agent cannot modify production tables and schema changes propagate only on publish. And we wire external systems through connectors or custom MCP servers, supported since December 2025, rather than through credentials pasted into application code.

Where the client is also a Databricks customer, we prefer the Databricks Apps deployment path so the application inherits Databricks authentication and governance rather than acquiring a second identity model. Where they are not, we keep the data access layer thin enough that moving later is a week, not a quarter.

What we do with this

We use Replit for speed and then apply the same review we would apply to any production service: deployment sizing, secret scope, data boundaries, and a rollback plan.

The guardrails we add on both

Neither platform removes the need for an exit. We build on interfaces that have institutional backing rather than vendor roadmaps. MCP moved to the Agentic AI Foundation, a Linux Foundation directed fund, on 9 December 2025, reporting over 97 million monthly SDK downloads and 10,000 active servers, with AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft and OpenAI as platinum members. Google donated Agent2Agent to the Linux Foundation on 23 June 2025, and A2A v1.0, the first production-ready stable release, landed in March 2026.

The inference interface is portable in practice. vLLM implements seven OpenAI-compatible endpoints; SGLang, hosted by the non-profit LMSYS, reports trillions of tokens per day across more than 400,000 GPUs behind OpenAI-compatible APIs; LiteLLM proxies 100+ providers behind a single OpenAI-compatible endpoint with routing, fallbacks and per-key budgets; and Bedrock exposes an OpenAI-SDK-compatible base URL directly. Azure has offered a v1 route usable with the stock OpenAI client and no api-version parameter since August 2025, which also serves non-OpenAI models. Google's Gemini compatibility layer works too, but is explicitly still in beta and silently ignores unsupported parameters — a correctness hazard rather than a feature gap.

Observability is the least settled piece, and we say so up front. The OpenTelemetry GenAI semantic conventions define an invoke_agent span with child chat and execute_tool spans plus attributes like gen_ai.request.model and token usage, and they are described as still under active development rather than stable. We instrument to them anyway — AgentCore Evaluations already accepts telemetry from any framework emitting those conventions or OpenInference — while treating attribute names as subject to change.

Cost levers are portable in shape but not in detail, and the sharp edges cost real money. Batch is 50% off on Anthropic, Bedrock and Azure, but Anthropic caps a batch at 100,000 requests or 256 MB and expires jobs after 24 hours, and Azure fails a job unless completion_window is exactly 24h. Bedrock prompt caching cannot be combined with the batch inference API, so the two levers do not compose, and cross-region inference may increase cache writes. Commitments are the least portable of all: Microsoft Foundry provisioned throughput reservations are not interchangeable across Global, Data Zone and Regional deployment types, unused reserved PTUs are lost each period, and cancellations count against a USD 50,000 rolling-12-month cap across all Azure reservations in the billing scope.

  • Every architectural choice goes into an append-only decision log. AWS frames a collection of ADRs as a hand-over artefact; Microsoft's Well-Architected guidance prescribes not editing accepted records, and instead linking superseded ones, with each entry carrying problem statement, options, outcome with trade-offs, confidence level and status.
  • Where Kubernetes is in play, we prefer platforms in the CNCF Certified Kubernetes AI Conformance Program, launched 11 November 2025 and grown from 18 to 31 certified platforms by 24 March 2026, with coverage extended to agentic workloads and sandboxing.
  • We assume multi-model from the start. Databricks says 63% of its customers already route across two or more model families, and Replit's own routing work claims 65% lower cost at equal quality.

What we do with this

The guardrails are the part clients keep when they change platforms: an open protocol layer, an OpenAI-compatible inference boundary, instrumented traces, and a decision log that explains why every one of those choices was made.

Sources

primary sources, checked on Sep 11, 2026
  1. 01Databricks Data + AI PlatformDatabricks · 2026-09
  2. 02Databricks Grows >80% YoY, Surpasses $7B Revenue Run-Rate, Scales Lakebase, Genie, and Unity AI GatewayDatabricks · 2026-08-13
  3. 03Agent Bricks: Data + AI Summit 2026Databricks · 2026-06-16
  4. 04Agent Bricks: The governed enterprise agent platformDatabricks · 2026-04-14
  5. 05Introducing Agent Bricks: Auto-Optimized Agents Using Your DataDatabricks · 2025-06-11
  6. 06Agent Bricks Supervisor Agent is Now GA: Orchestrate Enterprise AgentsDatabricks · 2026-02-10
  7. 07AI governance at Data + AI Summit 2026: What's new with Unity AI GatewayDatabricks · 2026-06-16
  8. 08Introducing Genie One, Genie Agents, and Genie OntologyDatabricks · 2026-06-16
  9. 09What's New in AI/BI - February 2026 RoundupDatabricks · 2026-02-11
  10. 10Databricks Lakebase is now Generally AvailableDatabricks · 2026-02-03
  11. 11Decoupled by Design: Billion-Scale AI SearchDatabricks · 2026-03-09
  12. 12Databricks AI Search (formerly Vector Search)Databricks · 2026-09-08
  13. 13MLflow 3 for GenAIDatabricks · 2026-06-30
  14. 14Announcing General Availability of Databricks AppsDatabricks · 2025-06-11
  15. 15Lakeflow: A new era of agentic data engineeringDatabricks · 2026-06-16
  16. 16Introducing Omnigent: A Meta-Harness to Combine, Control and Share Your AgentsDatabricks · 2026-06-13
  17. 17June 2026 Databricks platform release notesDatabricks · 2026-06
  18. 18Databricks-hosted foundation models available in Foundation Model APIsDatabricks · 2026-09
  19. 19Databricks and Anthropic Sign Landmark Deal to Bring Claude Models to the Data + AI PlatformDatabricks · 2025-03-26
  20. 20Databricks and OpenAI Launch Groundbreaking Partnership to Bring Frontier Intelligence to Enterprises with Databricks Agent BricksDatabricks · 2025-09-25
  21. 21Claude Fable 5 is now available on Databricks, fully governed through Unity AI GatewayDatabricks · 2026-06-09
  22. 22What's new with Unity Catalog at Data + AI Summit 2026Databricks · 2026-06-16
  23. 23Databricks Announces 2026 Data + AI Summit Keynote Lineup and ProgrammingDatabricks · 2026-06-15
  24. 24Databricks PricingDatabricks · 2026-09
  25. 25Mosaic AI Foundation Model Serving pricingDatabricks · 2026-09
  26. 26Announcing Databricks Lakebase Launch PartnersDatabricks · 2026-02-02
  27. 27Introducing Replit AgentReplit · 2024-09-16
  28. 28Introducing Replit Agent v2 in Early AccessReplit · 2025-02-25
  29. 29Introducing Agent 3: Our Most Autonomous Agent YetReplit · 2025-09-10
  30. 30Enabling Agent 3 to Self-Test at ScaleReplit · 2025-12-15
  31. 31Introducing Replit Agent 4: Built for CreativityReplit · 2026-03-11
  32. 32What's changed from Replit Agent 3 to Agent 4Replit · 2026-03-19
  33. 33Closing the loop: Evaluating and improving Replit Agent at scaleReplit · 2026-06-23
  34. 34Intelligent Model Routing on ReplitReplit · 2026-08-26
  35. 35Replit Introduces Free ModeReplit · 2026-08-18
  36. 36Exclusive: Replit taps OpenAI's low-cost Luna AI model for new 'Free Mode'Fortune · 2026-08-19
  37. 37Development and production databases — Replit DocsReplit · 2026-09
  38. 38About Deployments — Replit DocsReplit · 2026-09
  39. 39App Storage (Object Storage) — Replit DocsReplit · 2026-09
  40. 40Secrets — Replit DocsReplit · 2026-09
  41. 41Agent Integrations — Replit DocsReplit · 2026-09
  42. 422025: Replit in ReviewReplit · 2025-12
  43. 43Package Firewall: Blocking 8,000+ malicious packages dailyReplit · 2026-06-09
  44. 44Black-box pen tests on ReplitReplit · 2026-08-17
  45. 45AI-powered coding tool wiped out a software company's database in 'catastrophic failure'Fortune · 2025-07-23
  46. 46Ship Enterprise Data Apps Faster with Replit and DatabricksReplit · 2026-02-18
  47. 47Replit Enterprise, Now Self-ServeReplit · 2026-05-21
  48. 48Replit EnterpriseReplit · 2026-09
  49. 49Replit Closes $250 Million in Funding to Build on Customer MomentumPR Newswire · 2025-09-10
  50. 50The Future is Actually Very HumanReplit · 2026-03-11
  51. 51Inside Replit's path to $100M ARRGrowth Unhinged · 2025-03
  52. 52In a blow to Google Cloud, Replit partners with MicrosoftTechCrunch · 2025-07-08
  53. 53Introducing Composer 2Cursor · 2026-03-19
  54. 54Amazon Bedrock AgentCore is now generally availableAWS · 2025-10-13
  55. 55Release notes for Amazon Bedrock AgentCoreAWS · 2026-07
  56. 56Azure OpenAI in Microsoft Foundry Models v1 APIMicrosoft · 2026-05-13
  57. 57MCP joins the Agentic AI FoundationModel Context Protocol · 2025-12-09
  58. 58Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF)Linux Foundation · 2025-12-09
  59. 59Google Cloud donates A2A to Linux FoundationGoogle · 2025-06-23
  60. 60A year of open collaboration: Celebrating the anniversary of A2AGoogle Open Source · 2026-04-16
  61. 61CNCF Launches Certified Kubernetes AI Conformance ProgramCNCF · 2025-11-11
  62. 62CNCF Nearly Doubles Certified Kubernetes AI PlatformsCNCF · 2026-03-24
  63. 63Inside the LLM Call: GenAI Observability with OpenTelemetryOpenTelemetry · 2026
  64. 64OpenAI-Compatible Server — vLLM documentationvLLM · 2026
  65. 65Welcome to SGLangSGLang / LMSYS · 2026
  66. 66LiteLLM AI Gateway (LLM Proxy)LiteLLM · 2026
  67. 67gpt-oss-120b — Amazon Bedrock User GuideAWS · 2025-08-05
  68. 68OpenAI compatibility | Gemini APIGoogle · 2026
  69. 69Batch processing — Claude DocsAnthropic · 2026
  70. 70Prompt caching for faster model inference — Amazon BedrockAWS · 2026
  71. 71How to use global batch processing with Azure OpenAI in Microsoft Foundry ModelsMicrosoft · 2026-05-13
  72. 72Save costs with Microsoft Foundry Provisioned Throughput ReservationsMicrosoft · 2026-07-17
  73. 73Using architectural decision records to streamline decision-making during developmentAWS Prescriptive Guidance · 2026
  74. 74Maintain an architecture decision record (ADR)Microsoft · 2026-04-10

Keep reading

model historytimelinefrontier models 14 min

From GPT-1 to today: what actually changed in eight years of models

A dated walk through the model releases that changed how these systems are built, priced and deployed, from June 2018 to September 2026.

Eight years separate GPT-1's 117 million parameters and 512-token context from GPT-6 Astra's 1.05 million-token context. In between, four things changed the shape of the field: pre-training at scale, instruction tuning with human feedback, reinforcement learning for chain-of-thought reasoning, and a standard agent stack. Move the PaLM entry (5 April 2022) after the InstructGPT entry (4 March 2022) so the timeline actually runs in date order; leave this sentence as written.

you should know

Three recipe changes carried the field, not one: generative pre-training (GPT-1, June 2018), instruction tuning with human feedback (InstructGPT, March 2022), and reinforcement learning for chain-of-thought reasoning (o1, September 2024).

Sep 11, 2026Read
architecturetransformersinference 12 min

The architecture story: from the transformer to reasoning and agents

Nine ideas, grouped by what each one changed for people building products on top of these models.

Modern language models are the result of a sequence of separable ideas: the transformer, scaling laws, post-training, sparse experts, long context, inference efficiency, reinforcement learning for reasoning, multimodality and tool use, and finally protocols. This post walks the sequence in order, with dates and numbers from the cited sources, and states the practical consequence of each step for anyone shipping a product.

you should know

Parameter count no longer predicts serving cost. Sparse mixture-of-experts models activate a fraction of their weights per token: DeepSeek-V3 is 671B total but 37B active, and Mixtral 8x7B was 47B total with 13B active.

Sep 11, 2026Read

Let's build intelligent systems that drive growth

Tachyon is the engineering partner for teams that need AI in production, not in a deck. Start with a free 60-minute discovery call.