cross-posted from: https://futurology.today/post/2910566

Alibaba’s Qwen team just released QwQ-32B-Preview, a powerful new open-source AI reasoning model that can reason step-by-step through challenging problems and directly competes with OpenAI’s o1 series across benchmarks.

The details:

QwQ features a 32K context window, outperforming o1-mini and competing with o1-preview on key math and reasoning benchmarks.

The model was tested across several of the most challenging math and programming benchmarks, showing major advances in deep reasoning.

QwQ demonstrates ‘deep introspection,’ talking through problems step-by-step and questioning and examining its own answers to reason to a solution.

The Qwen team noted several issues in the Preview model, including getting stuck in reasoning loops, struggling with common sense, and language mixing.

Why it matters: Between QwQ and DeepSeek, open-source reasoning models are here — and Chinese firms are absolutely cooking with new models that nearly match the current top closed leaders. Has OpenAI’s moat dried up, or does the AI leader have something special up its sleeve before the end of the year?

  • SkaveRat
    link
    fedilink
    English
    arrow-up
    13
    arrow-down
    1
    ·
    1 day ago

    Love the step by step reasoning. Especially when you give it some weird stuff.

    I asked:

    How many strawberries can you fit into the most common sedan car from 2023, while also being able to drive it?

    The answer is too long to post directly:

    https://pastebin.com/mTiyQkmY

    • jcg@halubilo.social
      link
      fedilink
      English
      arrow-up
      10
      ·
      1 day ago

      I thought it’d drop the “just the trunk space” thing eventually but it reaffirms it towards the end

      But the question specifies that the car should still be drivable, which probably means that the rear seats need to be in place for passengers to sit.

      And the reasoning broke down, you don’t need passengers to drive a car. Pretty interesting reading it’s “thought” process with the little humanisms like “hmm” and “but wait!”