Skip to content

Bee Models

Specialist intelligence across every Bee tier.

Bee includes general text and conversation, but it is built for specialist work: reasoning, research, analysis, coding, tools, documents and governed workflows. Bee Cell through Bee Swarm are controlled model lines spanning 128K to 1M context. Every alias resolves to a governed production configuration, and an adapter can serve only after its held-out evaluation and release checks pass. Enclave is a private deployment contract. Bee Ignite is the active HEOSSI foundation research programme, but it is not a customer-selectable model and has no released checkpoint today.

Bee release system

Every Bee model is a controlled specialist-intelligence release.

Cell through Swarm are stable Bee product and API identities. Each release binds its serving configuration, specialist capabilities, safety policy, evaluations and immutable release evidence through the Bee Specialist Intelligence System.

Bee's hosted inference is operated by HEOSSI. Its serving configuration, governance, safety enforcement, routing, retrieval, tools, evaluations, and release controls are designed, implemented, and operated by HEOSSI through the Bee Specialist Intelligence System (BSIS).

Bee Ignite foundation programme

HEOSSI is developing the research, provenance and training control plane required for a future independently pretrained specialist foundation model through the active Bee Ignite foundation research programme. Ignite is not a transfer or adapter lane and has no released checkpoint today.

Explore BSIS

Model specifications and product evidence

Specialist intelligence, verified lane by lane.

Generated from the pinned release catalog and the live-verified capability source of truth. Last capability verification: 2026-07-14. Higher plans may route a task to an included lower lane; this table describes each named primary lane rather than the full set of services available in the product.

Inspect verification →
Bee model serving roles and live capability by tier
Bee tierArchitecture & servingContextText & reasoningNative imageNative videoTool-call APIStrict JSON1M
Bee Cell

Free, fast multimodal entry tier

4B dense · unified vision-language

Bee-hosted vLLM lane

128K native Live Live Live
Bee Brood

Cost-efficient multimodal reasoning

9B dense VL · thinking-enabled

Bee-hosted vLLM lane

256K native Live Live Live Live
Bee Comb

Structured multimodal workhorse

27B dense VL · FP8

Bee-hosted vLLM lane + 1M twin

256K native · 1M routed Live Live Live Live Live Live
Bee Buzz

Agent and builder workhorse

35B-A3B sparse MoE VL · FP8

Bee-hosted vLLM lane + 1M twin

256K native · 1M routed Live Live Live Live Live Live
Bee Hive

High-capability specialist intelligence

122B-A10B sparse MoE VL · NVFP4

Bee-hosted vLLM lane

256K native Live Live Live Live Live
Bee Swarm

Premium distributed intelligence

Versioned multi-lane routing policy

Governed routing fabric

1M native Live Routed Live

Core intelligence and workflows

General text and chat provide the baseline. Bee is differentiated by specialist reasoning, current-source research, coding and tools, document intelligence, durable workspaces, automation and governed long-context work.

General text & conversation

Live · all tiers · default

Every Bee model handles text generation, conversation, writing, summarisation and streaming responses. This is the baseline, not Bee's product category.

Specialist reasoning & analysis

Live · all tiers · tier-scaled

Bee is designed for technical, business, research and governed specialist workflows. Reasoning depth, context and assurance increase through the tiers.

Research & live web search

Live · all tiers · selectable

Web Search retrieves current sources with citations. Research mode plans multiple searches, reads across results and synthesizes a sourced answer.

Coding & tool use

Live · tier-aware

Coding assistance is available across the lineup. Structured tool calls are exposed through the lanes proven in the model matrix above.

Projects, memory & automation

Live · all tiers

Projects organise ongoing work, memory preserves useful context, and scheduled tasks run prompts and deliver notifications.

Document intelligence

Live · all tiers

Tenant-scoped retrieval, full-document injection when it fits, and citations ground specialist work in the user's own source material.

Context compaction

Live · automatic + manual

Every tier compacts at 75% of its context budget. Recent multimodal turns stay verbatim; older media is represented explicitly in the summary.

Supporting modalities and specialised compute

These services let specialist workflows accept and produce more than text. They extend Bee’s intelligence layer; they do not turn Bee into a media stack.

Speech input & output

Live · all tiers

Speech-to-text and text-to-speech are provider-routed Bee services available from Cell through Swarm.

Files & media attachments

Live · multimodal composer

The workspace accepts images, audio, video, PDFs, office files, code and text. Native video analysis remains limited to the lanes proven above.

Image generation

Live · plan-capped

Text-to-image generation is a prepaid, governed service separate from image understanding: a 3/day promotional safety ceiling on Cell and progressively higher funded-plan ceilings.

Video generation

Built · not active

The generation workflow is not sold as live until completed-job URL delivery passes production verification.

Quantum jobs

Opt-in · entitlement-gated

Simulator and OpenQuantum hardware jobs use a dedicated endpoint. Ordinary chat never silently routes through a QPU.

Bee model family

Stable model names across the workspace and OpenAI-compatible API. Availability follows your plan and current serving capacity.

Bee Cell

bee-cell · 128K context · base training-data cutoff not publicly disclosed

Current-source Web Search and Research available

Input

$0.30 / 1M

Output

$1.20 / 1M

Best for

  • Solo developers
  • Founders
  • Students and researchers
  • Developers exploring Bee
  • Fast everyday chat + coding

Capabilities

ChatVisionVoiceTool callingDocument RAGStreaming
Try Bee Cell

Bee Brood

bee-brood · 256K context · base training-data cutoff not publicly disclosed

Current-source Web Search and Research available

Input

$0.60 / 1M

Output

$2.40 / 1M

Best for

  • Solo developers running agentic flows
  • Founders needing thoughtful but cheap inference
  • Researchers iterating on prompts
  • Cost-conscious teams

Capabilities

ChatVisionVideo understandingVoiceTool callingThinking modeDocument RAGStreaming
Try Bee Brood

Bee Comb

bee-comb · 256K context · base training-data cutoff not publicly disclosed

Current-source Web Search and Research available

Input

$1.25 / 1M

Output

$5.00 / 1M

Best for

  • Software engineers
  • Security engineers
  • Technical founders
  • Dev teams
  • Researchers working locally

Capabilities

ChatVisionVideo understandingVoiceTool callingStructured output1M long-contextThinking modeDocument RAGStreaming
Try Bee Comb

Bee Buzz

bee-buzz · 256K context · base training-data cutoff not publicly disclosed

Current-source Web Search and Research available

Input

$2.50 / 1M

Output

$10.00 / 1M

Best for

  • Software engineers building agentic systems
  • Integration / platform engineering teams
  • Technical product teams
  • Mid-sized businesses with builder workflows

Capabilities

ChatVisionVideo understandingVoiceTool callingStructured output1M long-contextThinking modeDocument RAGStreaming
Try Bee Buzz

Bee Hive

Popular

bee-hive · 256K context · base training-data cutoff not publicly disclosed

Current-source Web Search and Research available

Input

$5.00 / 1M

Output

$20.00 / 1M

Best for

  • Startups
  • SaaS teams
  • Internal product / engineering teams
  • Small and mid-sized businesses
  • Teams building RAG / copilot workflows

Capabilities

ChatVisionVideo understandingVoiceTool callingStructured outputThinking modeDocument RAGStreaming
Try Bee Hive

Bee Swarm

bee-swarm · 1M context · base training-data cutoff not publicly disclosed

Current-source Web Search and Research available

Input

$10.00 / 1M

Output

$40.00 / 1M

Best for

  • Enterprise technical teams
  • Regulated businesses
  • Advanced engineering teams
  • Research labs
  • High-value analysis workflows

Capabilities

ChatVisionVoice1M long-contextThinking modeDocument RAGStreaming
Try Bee Swarm

Release evidence

A private training candidate is not a public model release. Bee promotes an adapter only after held-out evaluation, independent rubric comparison, and a reproducible snapshot are recorded. Validation criteria and release evidence are published at /trust.

Under the hood

Pinned releases, governed adapters, reproducible evidence.

The six public tiers are exposed as pinned, controlled Bee releases. Each tier uses a stable alias, preserves an immutable serving revision, and adds optional per-domain adapters only through an eval-gated release pipeline. That keeps multimodal, reasoning, tool, and long-context capability tied to reproducible evidence instead of a floating model label.

CorePinned Bee releaseimmutable serving revision
AttentionGrouped-Query + RoPElong-context friendly
Mixture of ExpertsSparse MoE routingcapacity without linear cost
ControlsThinking + efforttier-aware output budgets
RoutingTier-aware servingplan and capacity checked
EvidenceDated snapshotsrelease-id plus adapter digest
AdaptersPer-domain LoRAheld-out eval-gated
RetrievalBuilt-in RAGcite-back grounding

Capabilities

What the Bee platform can do.

Adaptive router

Difficulty-scored routing between local execution and frontier-teacher escalation. Scoring blends keyword complexity, query length, conversation depth, code/math detection, and a per-domain multiplier.

  • Difficulty-scored routing between a local lane and a frontier-teacher lane
  • Per-domain weighting for quantum, fintech, and cybersecurity queries
  • Self-verification (coherence + relevance + completeness) gates every response

source · Adaptive router

Evolution engine (architecture)

The evolution loop generates candidate neural modules — attention variants, SSM discretisations, compression codecs, memory protocols — runs them through a sandboxed eval, and only accepts winners. Architecture only; not currently running against the live Bee Cell deployment.

  • Population-based search over candidate modules per cycle
  • AST-based safety checks; forbidden imports/calls blacklist
  • Eval-gated acceptance, regression detection, automatic rollback

source · Evolution engine

Self-coding & self-healing (architecture)

When a request needs code, the self-coding module is designed to write it, execute in a sandbox, read the error, and iterate. It is enabled only for tiers and workflows that have cleared their release checks.

  • Self-coding: bounded iterate-and-execute loop in a sandbox
  • Self-healing: gradient norm monitor, loss-spike detection, NaN guard
  • Auto-tunes learning rate, checkpoints, rolls back to last good state
  • Algorithm invention, compression, crypto primitives, math proofs

source · Self-coding & self-healing

RAG pipeline

Document upload → chunk → embed → FAISS index → cite. End-to-end retrieval is built into every tier; pass document IDs in the chat request and Bee handles the rest.

  • Approximate-nearest-neighbour vector store
  • Sentence-embedding retrieval with cite-back grounding
  • Per-tenant isolation; shared index on Buzz and above

source · RAG pipeline

Quantum integration (opt-in)

Bounded real-QPU execution through OpenQuantum, with a local statevector simulator for development. Chat requests are never routed through quantum; Quantum Hardware Pack holders run explicit, metered jobs through the dedicated endpoint.

  • Provider: OpenQuantum
  • Reviewed targets: rigetti:cepheus-1-108q, iqm:emerald, iqm:garnet
  • Largest current target: 108 qubits
  • Local statevector simulator fallback (~28 qubits)
  • Five-credit per-job reserve; accepted quotes reconcile the customer allowance

source · Quantum integration

Domain adapters

LoRA adapters specialise the base model for low-cost fine-tuning per domain. Adapters are released through the governed-release pipeline once each clears the eval harness — every released adapter has a published validation record at /trust.

  • Domain-specific data collection · eval-harness-gated · governed release
  • General, technical, business, regulated, and research workflow domains
  • Restricted-domain (Tier-3) workflows include explicit acknowledgement and jurisdictional gates per /legal/acceptable-use
  • Specific LoRA configuration is contractual via the Enclave deployment manifest

source · Per-adapter validation records published at /trust

API compatibility

OpenAI Chat Completions — drop-in.

Bee speaks the same wire protocol as OpenAI's /chat/completions. Switch the base URL and the API key — your existing client code keeps working. Streaming is universal; tools, JSON mode, structured outputs, vision, and video are exposed on the tiers whose live model cards list them.

Deployment

Bee Enclave

Deployment mode

bee-enclave · contracted · knowledge boundary follows the selected governed lane

Customer-scoped Hive- and Swarm-class deployment designs can cover private, regulated, or sovereign requirements. Broader self-hosted and air-gapped operation remains roadmap.

Private Hive / Swarm workloadsAudit + compliance evidencePQC transport
Contact sales

Foundation research programme

Research work is active. Model serving is not. Bee Ignite becomes a model only after a HEOSSI-pretrained checkpoint clears the full BSIS release process.

Bee Ignite

Programme activeNo released checkpoint

bee-ignite · HEOSSI foundation programme · pretraining not started

HEOSSI-internal foundation-model programme inside BSIS. Bee Ignite requires a HEOSSI tokenizer, rights-qualified corpus, random-initialized architecture, reproducible distributed pretraining, specialist evaluation, shadow, canary and rollback evidence. No checkpoint is currently released or commercially available. Earlier transfer, adapter, and experimental serving lanes are retired; they do not establish independent foundation-model provenance.

Now

Programme controls and provenance

Next

Scaling, architecture and tokenizer

Then

Corpus and distributed pretraining

Release

Evaluation, shadow, canary and rollback

Programme workstreams

Architecture researchHEOSSI tokenizerRights-qualified pretrainingRelease-gate experiments