Skip to content

OSS Week 2026: AI Between Hype and Stance – Part 2

August 3rd, 2026

In Part 1 of our review series, we outlined the framework for this year’s OSS Week: our deliberate choice of values-aligned infrastructure, the providers we explicitly ruled out and why, and the pragmatic compromise we eventually settled on. We decided against providers like OpenAI or Google, opting instead for solar-powered cloud hosting with Infomaniak in Geneva and bringing in Anthropic as a reference.

One thing was clear: our focus was not on the technology alone, but on how we want to engage with AI – transparently, resource-consciously, and with digital sovereignty. The ethical, climate, and economic discussions produced a working framework for how we proceed.

With this foundation of stance and infrastructure in place, we spent that same week testing what was technically possible. Can open-weight models really keep up in everyday work? Where do they add value, where do they hit their limits? And what about the much-hyped vibe coding when put to the test under real-world conditions? Below, we share our concrete hands-on insights from that intensive week – along with the skepticism that stayed with us long after.

Open-Weight Models in Practice: Stronger Than Expected 

Can open-weight models keep up with the industry standard, or are they unusable in a professional context?

While the philosophical debates continued in earnest, working hands-on with open-weight models yielded some tangible findings. The performance and usability of current open-weight models is a mixed bag, but not hopeless. For the bulk of our non-coding tasks, Moonshot’s Kimi K2.6 model, for example, is perfectly adequate and does not lag significantly behind Claude. In coding tasks, too, open-weight models like Kimi or Qwen can definitely hold their own against Claude.

The Hidden Value: Tests and Documentation 

Where does AI deliver the greatest benefit in the day-to-day of software development?

When it comes to the core task of software development – writing production code – our experiments revealed clear limits. Neither proprietary nor open-weight models produced the elegance and technical coherence on screen that would have convinced us to adopt them in everyday project work.

Surprisingly, however, we found the greatest added value in two supporting workflows: documentation and writing tests. AI-assisted test generation proved especially valuable because the models operate free from our own cognitive biases. They regularly produce test cases that we ourselves, under time pressure or simply out of habit, would not have thought of. At the same time, the actual production code remains untouched, and the AI does not directly interfere with the architectural craft itself.

Digital Sovereignty and the Myth of Vibe Coding 

Do AI tools promise more than they can deliver for production code?

The fact that AI itself failed to convince us when it came to production code led us to scrutinize all the more critically the promise currently making the rounds: so-called vibe coding. The idea of quickly cobbling together something functional proved misleading in our experiments. The subsequent quality assurance and reworking can quickly consume more time than it saves. This raises a central question for us: Was the AI assisting us, or were we in fact assisting the AI by compensating for its errors and shortcomings?

Where Do We Go From Here?

These and further questions will continue to occupy us going forward: How do we ensure that a culture of prompting does not displace fundamental problem-solving competence? Is the open-source community ethos strengthened by the use of AI, or does human exchange dwindle? How can we remain productive when AI models are trained to maximize our interaction time? And how do we handle unmarked AI-generated code contributions without ultimately increasing the review burden on developers?

Our OSS Week demonstrated that using AI in a way consistent with our values doesn’t begin at the end of the process, but with smart deployment in the right places: where it supports without seducing, and complements human competence without replacing it. We’re keeping at it – and look forward to pursuing these questions together with the community and our partners.

Have you had similar experiences, or are you facing comparable trade-offs? We warmly invite you to join the conversation.