Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for NextBit

NextBit

Browse models provided by NextBit (Terms of Service)

10 models

Tokens processed on OpenRouter

  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Flash 0731DeepSeek V4 Flash 0731

    DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

    by deepseekJul 31, 20261.05M context$0.44/M input tokens$1.32/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Pro 0423DeepSeek V4 Pro 0423

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks. Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical

    by deepseekApr 24, 20261.05M context$1.72/M input tokens$3.45/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Flash 0423DeepSeek V4 Flash 0423

    DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

    by deepseekApr 24, 20261.05M context$0.14/M input tokens$0.28/M output tokens
  • Favicon for google
    Google: Gemma 4 26B A4B Gemma 4 26B A4B

    Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video (up to 60s at 1fps). Features a 256K token context window, native function calling, configurable thinking/reasoning mode, and structured output support. Released under Apache 2.0.

    by googleApr 3, 2026262K context$0.10/M input tokens$0.40/M output tokens
  • Favicon for qwen
    Qwen: Qwen3 14BQwen3 14B

    Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, programming, and logical inference, and a "non-thinking" mode for general-purpose conversation. The model is fine-tuned for instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.

    by qwenApr 28, 2025132K context$0.10/M input tokens$0.22/M output tokens
  • Favicon for sao10k
    Sao10K: Llama 3.3 Euryale 70BLlama 3.3 Euryale 70B

    Euryale L3.3 70B is a model focused on creative roleplay from Sao10k. It is the successor of Euryale L3 70B v2.2.

    by sao10kDec 18, 20248K context$0.65/M input tokens$0.75/M output tokens
  • Favicon for thedrummer
    TheDrummer: UnslopNemo 12BUnslopNemo 12B

    UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

    by thedrummerNov 8, 202432K context$0.40/M input tokens$0.40/M output tokens
  • Favicon for google
    Google: Gemma 2 27BGemma 2 27B

    Gemma 2 27B by Google is an open model built from the same research and technology used to create the Gemini models. Gemma models are well-suited for a variety of text generation tasks, including question answering, summarization, and reasoning. See the launch announcement for more details. Usage of Gemma is subject to Google's Gemma Terms of Use.

    by googleJul 13, 20248K context$0.65/M input tokens$0.65/M output tokens
  • Favicon for undi95
    ReMM SLERP 13BReMM SLERP 13B

    A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge

    by undi95Jul 22, 20234K context$0.45/M input tokens$0.65/M output tokens
  • Favicon for gryphe
    MythoMax 13BMythoMax 13B

    One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge

    by grypheJul 2, 20234K context$0.06/M input tokens$0.06/M output tokens