studies show a clear trend – output is up (more code, more commits, bigger diffs), but outcomes don’t reflect that trend. If anything, the average team is taking longer to ship worse software

  • Feyd@programming.dev
    link
    fedilink
    arrow-up
    20
    ·
    2 hours ago

    output is up (more code, more commits, bigger diffs)

    We’ve known measuring output by lines of code is counterproductive for a long time.

  • ArseAssassin@sopuli.xyz
    link
    fedilink
    arrow-up
    11
    ·
    2 hours ago

    One study found a significant correlation between confidence in AI output and belief in the paranormal.

    💀 💀 💀

  • ell1e@leminal.space
    link
    fedilink
    English
    arrow-up
    8
    ·
    edit-2
    2 hours ago

    This doesn’t seem to cover there is also no LLM that doesn’t plagiarize, or where the training data appears to be compatible with such behavior (e.g. CC0). Now I don’t know what that means legally, but morally it seems to be tossing away other project’s licensing and I think for FOSS as a whole that’s no good.

    Also something worth reiterating: https://machinelearning.apple.com/research/illusion-of-thinking LLMs apparently can’t do basic logical reasoning. Even a junior coder can do that. I’m always surprised anybody would let LLMs near their code, at all.

  • MagicShel@lemmy.zip
    link
    fedilink
    English
    arrow-up
    5
    ·
    2 hours ago

    This aligns with my experience, largely. Of course it’s still my job to maximize LLM effectiveness within my organization. Which is a delicate balancing act to protect my teams from overeager executive leadership looking for huge gains.

    My own summary is that AI can be an accelerator, but the harder you lean into it, the worse outcomes will be. No matter how much code is written, you still need actual human minds to understand it and they can only handle so much volume before getting overwhelmed.

    Also, if AI gives you 20% productivity gains, but that 20% goes into playing with AI trying to get more, you haven’t really gained anything. Usage needs to be standardized rather than developers constantly negotiating with AI trying to coax out better outcomes.

  • faltryka@lemmy.world
    link
    fedilink
    arrow-up
    6
    ·
    3 hours ago

    Some of this does not line up with my lived experience pretty starkly.

    Repo level markdown files with architectural guidance not working for example… I’ve found that works quite well.

    Not perfectly well, but llms are designed specifically NOT to be perfect deterministic executioners. Still though, pretty well.

    I have seen that in a jr engineers hands llms get to bad outcomes fast, and unintuitively (to leaders…) usage of llms in coding does not provide a path for a he engineer to upskill into a sr engineer. A sr engineer with llms though is almost always radically augmented regarding their output speed on task completion.

    • MagicShel@lemmy.zip
      link
      fedilink
      English
      arrow-up
      3
      ·
      2 hours ago

      I agree with your last paragraph. We had about 6 weeks of unlimited AI spend before the costs reached executive leadership, and in that time I saw the least experienced developers spend the most with the least to show for it.

      But I will say that another factor is thinking that if you get 10% gains from a little AI, then a lot of AI will get you 100%.

      But I find the article is right about repo-wide docs. At least on their own. I find having small markdowns (often in the form of skills/commands), focused on specific tasks reduces spend (especially when your execution agent is a low cost model, leaving the reasoning to dedicated agents) and gives better outcomes. Loading massive docs into every task reduces the attention to the task at hand and often confuses AI as the reasoning part of the model becomes overwhelmed and starts inferring wrong things confidently.

      I suppose it heavily depends on the scale of the repo though. A large microservice with multiple upstream services it needs to call spends a lot tokens on API which is unnecessary for most tasks. And then it decides to use the wrong one… I have stories lol.

    • WhatAmLemmy@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      14 minutes ago

      It’s obvious if you actually use software beyond the average literacy of a talking chimp. I use hundreds of apps across iphone, mac, and linux os’s. Literally none of them have noticeably increased in quality, stability, or feature-set beyond their average between 1-5 years ago.

      Mac and iphone appear to have more bugs and shittier quality control than at any other point in the last decade.

      Quality software is getting harder and harder to find thanks to all the slop abandonware being promoted by slop content and slop SEO on slop driven search engines.

      I notice far more idiocracy-grade errors in digital content, cx, business processes than ever before.

      Auto-generated subtitles are great for content that was never going to receive human attention, but they’re clearly being used to replace humans. At least once a week I notice a major contextual error that completely alters the perception of the line/scene, and there’s no way to submit corrections. Art, culture, and knowledge are being actively corrupted and bastardized.

  • OpenStars@discuss.online
    link
    fedilink
    English
    arrow-up
    2
    ·
    3 hours ago

    Hehehehe

    LLMs struggle with negation. Telling them not to do something can often have the same effect as telling them to do it.

    The future looks… unreliable.

    Models may get more powerful, but not significantly more reliable. This it folks – work with what you’ve got!

    We all know that true AGIs becoming smarter than humans seems inevitable, but that could be like a hundred years from now, if ever. What’s unclear is what will happen two years from now, involving matters having little to do with the technology & what it is capable of and instead more to do with the economy and what jobs will be available then.

    img

    • chuckleslord@lemmy.world
      link
      fedilink
      arrow-up
      11
      ·
      2 hours ago

      AGI isn’t possible with current or near tech. Anyone who says otherwise is huffing paint or selling AI crap.

      • Initech_vs_Initrode@lemmy.zip
        link
        fedilink
        arrow-up
        1
        ·
        28 minutes ago

        Wow, what a solid argument you’ve made. It definitely doesn’t reek of someone trying to hype themselves up in the face of an unknown threat.

        • chuckleslord@lemmy.world
          link
          fedilink
          arrow-up
          1
          ·
          49 seconds ago

          We invented a next token predictor. That isn’t intelligence nor is it on the path to it, either. The word rocks aren’t any closer to real intellect than the math rocks were. You just think they are cause words are the things that humans use to communicate across time and space.

          The only ones spouting nonsense that it is are those whose business models require AGI to be achievable within the next decade. But, ya know, wish in one hand and all that.

      • OpenStars@discuss.online
        link
        fedilink
        English
        arrow-up
        2
        ·
        49 minutes ago

        I literally said “if ever”, and also “seems” rather than “is”. I also never so much as implied current tech, with my comment about a hundred years from now.

        You are reacting against what I never said.

        Though I choose to upvote your comment anyway, since at least you said it rather than simply assumed it and moved on.

    • Blurntout@lemmy.ca
      link
      fedilink
      arrow-up
      3
      ·
      3 hours ago

      15 - 30 years of pain followed by the end of life as we know it those who adapt may thrive but will continue to be exploited.