AI hasn't replaced software developers, and the evidence so far suggests it won't any time soon. What it has done is change how code gets written, reviewed and tested. For anyone paying for software, the practical question isn't "does my vendor use AI?" (almost all of them do) but "does their use of AI make my product faster to build and safer to run?" The honest answer depends far more on the team's engineering discipline than on the tools themselves.
Key takeaways
- AI tools are close to universal: 90% of software professionals used AI at work in Google's 2025 DORA research, and 84% of developers use or plan to use AI tools (Stack Overflow, 2025).
- Controlled studies show speed-ups of roughly 21-56% on well-defined tasks (Google, 2024: about 21%; GitHub/Microsoft, 2023: 55.8%), but a 2025 METR trial found experienced developers were 19% *slower* on real work in codebases they knew well.
- AI increases throughput, but DORA has found it continues to have a negative relationship with delivery stability (2025). Testing, review and small releases decide whether speed turns into a better product.
- Generated code carries real security risk: 45% of samples in Veracode's 2025 study introduced OWASP Top 10 vulnerabilities.
- Ask vendors who is accountable for AI-generated code, how it's reviewed and tested, and where your data goes.
How widely is AI used in software development today?
AI assistance is now the default, not the exception. Google's 2025 DORA report found adoption among software professionals had reached 90%, with a median of two hours a day spent working with AI tools. The Stack Overflow Developer Survey 2025 reports that 51% of professional developers use AI tools daily.
Usage and confidence are not the same thing, though. In the same Stack Overflow survey, 45.7% of respondents said they distrust the accuracy of AI output, compared with about a third who trust it, and positive sentiment fell to around 60% from more than 70% in 2023 and 2024. The most common frustration, cited by 66%, was "AI solutions that are almost right, but not quite". AI produces plausible code quickly; someone still has to decide whether it's correct.
Where AI genuinely helps
AI is most useful when the task is well defined, the output can be checked automatically, and a mistake is cheap to catch. Those conditions describe a surprising amount of day-to-day engineering work.
Code completion and boilerplate
Autocomplete-style assistants are the most mature use. They're good at the repetitive parts of a codebase: data models, API handlers, form validation, configuration files. A 2023 controlled experiment by GitHub and Microsoft researchers found developers using Copilot completed a task 55.8% faster than a control group. The task was writing an HTTP server in JavaScript from scratch, which is exactly the kind of bounded, familiar problem these tools handle well.
Test generation
Writing unit tests is valuable, tedious and often skipped under deadline pressure. AI can draft test cases, edge cases and fixtures quickly, and the results are easy to verify: the tests either run and exercise the code or they don't. The catch is that generated tests can confirm what the code does rather than what it should do, so an engineer still needs to check that the assertions reflect the requirements.
Code review support
AI reviewers can flag likely bugs, inconsistent naming, missing error handling and unclear logic before a human reviewer looks at a pull request. The 2024 DORA report associated a 25% increase in AI adoption with a 3.1% increase in code review speed and a 3.4% increase in code quality. These are modest gains: a useful first pass that frees human reviewers to focus on design and business logic.
Documentation
Documentation showed the largest improvement in the same DORA analysis: a 7.5% increase in documentation quality for a 25% increase in AI adoption. AI is good at summarising what a function or module does, drafting API references and turning commit histories into readable change notes. For clients, better documentation means less dependency on any single developer and easier handovers.
Understanding and migrating legacy code
Older systems are often poorly documented and written in frameworks few people still know well. AI can explain unfamiliar code, map dependencies and draft translations from one language or framework to another. This can shorten the discovery phase of a modernisation project considerably. It does not remove the need for thorough regression testing, because a "translation" that looks right can still change behaviour in subtle ways.
Where AI doesn't help (yet)
The limits matter as much as the strengths, and they're mostly about judgement.
| Area | What AI does well | Where human judgement is still required |
|---|---|---|
| Requirements | Summarising notes, drafting user stories | Understanding the business, resolving conflicting stakeholder needs, deciding what not to build |
| Architecture | Explaining patterns, comparing options | Choosing trade-offs for your scale, budget, team and compliance obligations |
| Security | Spotting common vulnerability patterns | Threat modelling, access-control design, judging real-world risk |
| Code | Drafting functions, refactors, boilerplate | Verifying correctness, maintainability and fit with the wider system |
| Testing | Generating unit tests and test data | Deciding what must be tested and whether tests reflect the requirements |
| Delivery | Writing pipeline config and scripts | Release strategy, rollback planning, incident response |
Requirements and product decisions. An AI tool can't sit in a workshop and notice that two departments mean different things by "approval". Getting requirements right is still the single biggest factor in whether a project stays on budget, and it's a human conversation.
Architecture. Choosing between a monolith and services, how to structure a multi-tenant data model, or where to draw security boundaries depends on context the tool doesn't have: your growth plans, your regulators, your operations team.
Security judgement. AI can repeat common patterns, including insecure ones. Deciding what an attacker would target in your system requires threat modelling by people who understand it.
What the productivity evidence actually says
The research is genuinely mixed, and buyers should be wary of anyone who quotes only one side.
| Study | Year | Who | Result |
|---|---|---|---|
| Peng et al. (GitHub/Microsoft) | 2023 | Developers given one bounded task | 55.8% faster with Copilot |
| Cui et al. field experiments | 2024 | 4,867 developers at Microsoft, Accenture and a Fortune 100 firm | 26.08% more completed tasks; larger gains for less experienced developers |
| Paradis et al. (Google RCT) | 2024 | 96 Google engineers on an enterprise-grade task | About 21% less time, with a wide confidence interval |
| METR | 2025 | 16 experienced open-source developers, 246 real issues in their own repositories | 19% slower with AI tools |
How can both be true? The studies measure different things. The positive results come mostly from bounded tasks or broad populations that include many less experienced developers. The METR study looked at highly experienced developers working on large, mature codebases they already knew intimately, where the cost of reviewing and correcting AI output can outweigh the time saved.
The most telling METR finding is about perception. Developers expected AI to speed them up by 24%, and even after the study they believed it had sped them up by 20%, despite the measured slowdown. Self-reported productivity gains, including those in vendor marketing, should be treated with caution.
The organisational picture adds another layer. The 2024 DORA report found that a 25% increase in AI adoption was associated with a 1.5% decrease in delivery throughput and a 7.2% reduction in delivery stability. By 2025, DORA reported a positive relationship between AI and throughput, but the negative relationship with stability remained. DORA's own summary is that AI "doesn't fix a team; it amplifies what's already there."
In other words, AI amplifies developers rather than replacing them. A disciplined team with good tests, small releases and fast feedback gets faster. A team without those foundations ships more changes and more incidents.
The risks buyers should understand
Security of generated code
Veracode tested more than 100 large language models and found that 45% of generated code samples failed security tests and introduced OWASP Top 10 vulnerabilities. Java had a 72% failure rate, and models failed to defend against cross-site scripting in 86% of relevant samples. Just as important, security performance stayed flat as models got larger and newer. Better models write more working code, not necessarily safer code.
The practical response is the same as for human-written code, applied without exception: automated security scanning in the pipeline, peer review, and periodic independent security audits for systems that handle sensitive data.
Data leakage
When developers paste code, logs or customer records into an AI tool, that data may leave your control, depending on the tool's terms and configuration. OWASP lists sensitive information disclosure as the second-highest risk in its 2025 Top 10 for LLM applications. Your vendor should be able to tell you which tools they use, whether those tools retain or train on inputs, and how secrets and personal data are kept out of prompts.
IP and licensing
AI models are trained on large bodies of existing code, and the legal position on ownership and licence obligations for generated output is still developing in several jurisdictions. We won't offer legal advice here, but it's reasonable to expect your contract to state clearly that you own the delivered code, and to ask whether the vendor uses tool settings that filter suggestions matching public code.
Over-reliance and skills drift
If a team accepts suggestions without understanding them, the codebase becomes harder to maintain over time. The Stack Overflow finding that 45.2% of developers say debugging AI-generated code is more time-consuming is a warning sign here.
Governance: what good AI use looks like
Good AI governance in a development team is not complicated, but it has to be deliberate. DORA's 2025 recommendations start with clarifying AI policies, connecting AI to internal context, prioritising foundational practices and strengthening safety nets. In practice, that usually means:
- A named human owner for every change. AI can draft; an engineer approves, merges and answers for it.
- An approved tool list with data-retention settings checked, and a rule against putting secrets or personal data into prompts.
- Automated testing and security scanning on every pull request, run through a properly configured CI/CD pipeline.
- Small, frequent releases with fast rollback, which is where mature DevOps practices earn their keep: they contain the instability that faster change can bring.
- Extra scrutiny for high-risk areas such as authentication, payments and personal data.
This is the approach we follow at VulcanTech. Our teams use AI tooling for code review, test generation and documentation to ship faster, with a senior engineer accountable for the code that reaches production. The tools speed up the work; responsibility stays with people.
What to ask a vendor about their AI use
If you're comparing development partners, these questions separate thoughtful AI use from marketing. (For the broader selection process, see our guide to choosing a software development partner.)
- Which AI tools do your developers use, and for what? A clear, specific answer is a good sign. "We use AI for everything" is not.
- Who reviews AI-generated code, and how senior are they? Look for mandatory human review by experienced engineers.
- What automated tests and security scans run before code is merged? Ask to see an example pipeline.
- Where does our code and data go when your team uses AI tools? Retention, training opt-outs and region should all have answers.
- Who owns the code you deliver? It should be you, in writing.
- How do you measure whether AI is actually helping? Delivery metrics such as lead time, change failure rate and recovery time are more credible than "we're 50% faster".
- How will AI affect our cost and timeline? Expect some savings on routine work, but be sceptical of dramatic discounts that imply review and testing are being cut.
Frequently asked questions
Will AI make my software project cheaper?
It can reduce effort on routine work such as boilerplate, tests and documentation, which may shorten timelines. But requirements, architecture, integration and quality assurance still make up much of a project's cost, and the research shows gains vary widely by task and team. Treat large promised savings with caution.
Is AI-generated code less secure than human-written code?
It can be. Veracode found 45% of AI-generated samples introduced OWASP Top 10 vulnerabilities, and security didn't improve with newer models. The fix is the same discipline applied to all code: review, automated scanning and security testing before release.
Should I stop my vendor from using AI tools?
Usually not. Banning AI outright may cost you real efficiency gains. It's more useful to require clear governance: approved tools, data protections, mandatory human review and automated testing.
Can AI help modernise our legacy system?
Yes, particularly in understanding undocumented code and drafting migrations. It's a strong accelerator for discovery, but every migrated component still needs thorough regression testing by engineers who understand the original behaviour.
Does VulcanTech build AI features as well as using AI tools?
Yes. Separately from using AI in our development process, we build AI and machine learning solutions for clients, such as the AI and GIS air-quality dashboard we delivered for EPCCD, Government of the Punjab.
Conclusion
AI is changing software development in real, measurable ways: faster drafting, better documentation, quicker reviews. It is not changing the fundamentals. Good requirements, sound architecture, thorough testing and accountable engineers still decide whether a project succeeds. The evidence shows AI amplifies whatever process it lands in, so the most important thing to evaluate in a vendor is the process, not the tools.
If you're planning a web application or another software project and want to understand how AI-assisted development would affect your cost, timeline and risk, book a free 30-minute discovery session with our team.

