30 DAYS AGO • 6 MIN READ

AI Is Turning Into A Private Members Club. Is Open Source The Future?

profile

Connect The Dots

Morning All,

Anthropic's Fable lasted three days, Mythos never went public at all, and now OpenAI's latest frontier model is invite-only.

Last week OpenAI previewed GPT-5.6 in three tiers and at the request of the U.S. government, none of the three is going straight to the public.

The pattern is hard to miss.

The most capable models are starting to arrive late, locked, and approved one customer at a time.

Access to the greatest intelligence humanity has ever had is turning into something you need to apply for. If your face doesn't fit, or maybe more importantly if your views don't, then you'll have to get in the queue.

Or maybe they won't let you in the club at all.

So how do you bring the party to you? How do you make sure you're always having a great time with AI regardless of whether you're on the guest list or not?

For every individual worker and business decision maker, the answer to that question is becoming more important by the day


Choosing the Right LLM: Frontier Cloud vs Open-Source Cloud vs Local Open-Source (as of June 2026)

TL;DR

  • There is no single best LLM - the right choice depends on the task AND on three competing priorities:
    • Capability (frontier cloud wins)
    • Control/privacy (local open-source wins)
    • Value (open-source cloud wins).

For most non-technical knowledge workers, a frontier cloud model (Claude Opus 4.8, GPT-5.5, or Gemini 3.1 Pro) is the sensible default. Open-source cloud models (Kimi K2.6, DeepSeek V4, Qwen 3.7, GLM-5.2) are the smart value/customisation option. Small local models (Gemma 4, Qwen3.5, Phi-4, gpt-oss-20b) are the private, offline, zero-cost options that are now "good enough" for everyday text work, but still weak on hard reasoning.

  • The difference between the top models is so small now that cost, privacy, ecosystem fit, and the specific task matter more than raw "intelligence" rankings.
  • Local models hosted on a typical business laptop (e.g Lenovo Thinkpad with 16-32GB RAM, no dedicated GPU) can now do everyday work at roughly 80-90% of ChatGPT et al quality. Meaning they are now genuinely useful for text tasks such as brainstorming, 1st draft writing, summarising, and simple coding. But, they fall apart on complex multi-step reasoning, and they are not realistically viable for image, video, or design work on integrated graphics.

Important things to note:

  • Benchmarks are a compass, not a map. The gap between these top models is legitimately razor thin. Ranking no.1 on a leader board doesn't guarantee better results on your specific documents. Have the mindset of a scientist and run tests on your own real tasks.
  • The landscape is moving extremely quickly. All model names, versions, prices and rankings here are accurate up until June 2026 and will date quickly. GPT-5.6 was announced last week as a gated limited preview, Gemini 3.5 Pro is coming soon, and Kimi K2.7 just launched and is already looking like it's a step above K2.6.
  • Open video models are heavy. Open-weight video (Wan 2.7, LTX-2.3, HunyuanVideo 1.5) are good but you'll need some serious GPU's to run them.
  • Local model quality are workload-dependent. "80-90% of frontier quality" holds for everyday drafting, summarising and Q&A. The gap is much larger for complex reasoning, long-context work, and polished creative writing. Reasoning models across all tiers also hallucinate more than simpler models.
  • Privacy defaults change regularly. Whatever tool you use, the way to opt-out of having your data used to train future models changes on a regular basis. Make sure you confirm current settings in your app settings, before trusting any cloud tool with your sensitive material.

That being said, if you want the best AI setup, here are your best bets as of June 2026:

COMPARISON TABLE

Frontier context (June 2026): GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro and Grok 4.3 are the four flagships; and the gap between them is the lowest ever recorded. Where a stronger model (e.g. Fable 5) is suspended, the table uses the best available substitute.

Local constraint: "Local" = runs sensibly on a typical business laptop such as a current Lenovo ThinkPad.

Key Findings

  1. The frontier lineup as of June 2026 is OpenAI GPT-5.5 (released 23 April 2026, now the default ChatGPT model), Anthropic Claude Opus 4.8 (28 May 2026), Google Gemini 3.1 Pro (19 February 2026), and xAI Grok 4.3 (April 2026). Anthropic's even stronger Claude Fable 5 and Mythos 5 launched in June but were suspended on 12 June 2026 under a US Commerce Department export-control directive (expected to return for US users around 1 July 2026), so Opus 4.8 is the practical top Claude model right now.
  2. Open-weight (open-source) models have closed most of the gap. Kimi K2.6, DeepSeek V4, Qwen 3.5/3.7, GLM-5.2, MiniMax M3, and Llama 4 are almost as good as the closed frontier models on coding, reasoning and writing, and are 10-30x cheaper.
  3. Local models that fit on a business laptop are real but limited. Gemma 4 (E4B/12B), Qwen3.5 9B, Phi-4 14B, gpt-oss-20b and qwen2.5-coder run well on the latest Lenovo Thinkpad. Image generation needs a hefty GPU and video generation is effectively impossible on a regular laptop.

KEY CRITERIA/TERMINOLOGY

  • Accuracy: How often it is right, especially on hard reasoning, maths and factual questions.
  • Speed: How fast it responds.
  • Cost: Price per use: per-token API fees, monthly subscription, or one-off hardware.
  • Availability/access: What you need to use it (account, internet, hardware) and how exposed you are to outages and limits.
  • Privacy: Whether your prompts and files leave your device, and whether they train the model.
  • Data security: Breach exposure and who controls the data.
  • Ease of use/setup: How much technical effort before you get value.
  • Context window: How much text/data it can consider in one go.
  • Reliability/consistency: How dependably it follows instructions and formats.
  • Multimodal capability: Whether it handles images, audio and video, not just text.
  • Offline capability: Whether it works with no internet.
  • Customisation/fine-tuning: Whether you can adapt it to your own data.
  • Ecosystem/integrations: How well it plugs into the tools you already use.
  • Vendor lock-in: How hard it is to switch later.
  • Compliance/regulatory: Is it fit for regulated data and legal obligations.

Recommendations

Pick your default by what you value most.

  • If you mostly want the best answer with the least fuss and you're not handling sensitive data: use a frontier cloud model.
    • Claude Opus 4.8 is the safest all-rounder (writing, coding, analysis)
    • Gemini 3.1 Pro is the best value and best for research/long documents
    • GPT-5.5 is the broadest ecosystem
    • Grok 4.3 for live/social data and looser guardrails.

A single £20/month subscription on any of them covers most knowledge workers.

  • If you handle confidential or regulated data, work offline, or want zero per-use cost: install Ollama or LM Studio and run Gemma 4 12B or Qwen3.5 9B on your own machine. You'll have to accept that it is slower and weaker on hard reasoning.
  • If you (or your company's IT team) want cost control, customisation, or no vendor lock-in at scale: use open-weight cloud models (DeepSeek V4, Qwen 3.7, GLM-5.2, Kimi K2.6) via a hosted API.

Match the model to the task using the table above rather than loyalty to one brand. The single biggest mistake is forcing one tool to do everything.

For creative/visual work, stay in the cloud. Image, video and design are where local options are weakest. On a business laptop, don't attempt local image or video generation; use Nano Banana Pro / ChatGPT Images / Veo 3.1, or open-weight FLUX.2 / Wan 2.7 via a hosted service.


Thresholds that would change these recommendations:

  • Buy a 32GB+ machine or a GPU: At this point it will set you back thousands, and I don't know about you but...the way my bank account set up... However, if you need local image generation, longer context, or faster local responses...there is no other path, 16GB is the floor.
  • Switch to open-weight cloud: If your API spend on a frontier model gets to big enough volume this becomes the smart move. Open-weight is 10-30x cheaper per token.
  • Update your thinking regularly: The frontier reshuffles fast (three flagship releases in spring 2026 alone), and a model that led last month may not be leading this month. Build habits and workflows that are model-agnostic.
  • Watch availability shocks: The Fable 5 suspension shows frontier access can vanish overnight; keep a fallback model in mind just in case you get kicked out the club with no warning.

If you remember nothing else: There is no single best model and no single best model provider. Given that fact, locking yourself into an expensive contract with one company is not the best use of your personal or your company's money. Yes, that includes those companies that blindly signed up to Microsoft Copilot contracts, just because of sharepoint access. Building workflows that use multiple different models for different tasks is the most efficient, cost effective and productive way to use AI. If you can make yourself more productive whilst saving your company time and money, you become an extremely more valuable employee.


Benchmarks referenced : GPQA Diamond and Humanity's Last Exam (reasoning), SWE-bench Verified and SWE-bench Pro (coding), EQ-Bench Creative Writing and LMArena Text (writing/chat), Artificial Analysis Intelligence Index (overall), ARC-AGI-2 and FrontierMath (hard reasoning/maths), plus the LMArena Text-to-Image and Text-to-Video boards (creative). Sources include official model cards (Google DeepMind, xAI, DeepSeek, Anthropic), Artificial Analysis, LMArena/Arena.ai, llm-stats, vals.ai, EVY, Thunder Compute and Pinggy.

Connect The Dots