AI testing tools: Easy to buy but hard to trust?
Artificial intelligence is rapidly becoming embedded across software testing, promising everything from automated test generation and bug summarisation to visual validation and intelligent coverage analysis.
Yet as financial institutions accelerate AI adoption across software delivery, a growing number of quality assurance leaders are warning that the biggest challenge may no longer be whether AI can generate more tests, but whether those tests can actually be trusted.
For banks operating under increasingly stringent operational resilience, governance and auditability requirements, the distinction matters.
AI may dramatically increase testing output, but volume alone does not necessarily translate into confidence. In regulated financial environments, where missed edge cases can trigger compliance failures, operational outages or customer harm, QA teams are increasingly being asked to validate not only software, but also the AI tools helping to test it.
Consultant Margarita Simonova, the founder of consultancy firm ILoveMyQA, argues that organisations risk equating AI adoption with improved software quality, when the reality is considerably more nuanced.
“Buying an AI testing tool is not the same as building a stronger QA process,” Simonova said.
“A tool can create more tests, but it does not automatically create more confidence. It can make the dashboard look more active, but it does not always tell you whether the right risks were tested.”
“Buying an AI testing tool is not the same as building a stronger QA process.”
– Margarita Simonova
The warning reflects a broader shift taking place across banking QA. Financial institutions increasingly rely on AI to generate test cases, recommend automation, analyse logs and identify anomalies, but testing leaders are simultaneously placing greater emphasis on explainability, reproducibility and business context rather than sheer automation.
Simonova is careful to stress that AI itself is not the problem. On the contrary, she believes it has the potential to make QA teams significantly more effective when used as an intelligent assistant rather than an autonomous decision-maker.
“I do not think QA teams should avoid AI testing tools. If used well, they can be very helpful,” she stressed.
“AI can help testers create first drafts of test cases. It can suggest scenarios that a team may not think about right away. It can summarize long bug reports or logs. It can help identify duplicate issues.”
Creativity in testing
That view is consistent with comments Simonova made earlier, where she described AI as unlocking a more creative approach to software testing rather than simply accelerating existing automation.
“Efficiency is not the only advantage of AI, it is also about creativity,” Simonova argued.
“Traditional automated testing has not necessarily been viewed as a creative endeavour. It is typically somewhat rigid, with predefined paths and predictable outcomes. But when AI is inserted into the process, a type of dynamic creativity is unlocked.”
She believes one of AI’s greatest strengths lies in generating richer testing scenarios based on user behaviour, allowing teams to explore edge cases that conventional automation might never consider.
“One of the main areas where AI is useful is in generating test scenarios,” Simonova continued. “AI can analyse vast amounts of user behaviour data to generate creative, real-world test scenarios that human testers might not have even considered.”
Yet those capabilities should not be mistaken for judgement. According to Simonova, the greatest danger is not catastrophic AI failure, but the gradual development of misplaced confidence.
“The biggest risk with AI testing tools is not that they will fail completely,” she pointed out. “The bigger risk is that they will look useful enough to trust too quickly.”
She illustrated the problem with a practical example. An AI system may successfully generate hundreds of tests for a customer journey, covering all the obvious user actions.
But the genuine production risk may lie in highly specific business logic, such as regulatory requirements, payment edge cases, browser-specific behaviour or reporting errors that only experienced testers recognise as commercially significant.
“A tool may generate 100 test cases from a requirement. But are they the right 100? Do they cover the actual customer journey? Do they understand the business rules? Do they know which flow affects revenue, compliance, user trust or support volume?”
“The biggest risk with AI testing tools is that they will look useful enough to trust too quickly.”
– Margarita Simonova
For banks, those questions extend directly into regulatory obligations around resilience, traceability and software governance. Modern financial platforms combine payment systems, customer onboarding, fraud controls, cloud infrastructure and AI-driven services, making business context just as important as technical correctness.
Simonova argued that organisations should stop evaluating AI testing platforms purely on productivity metrics.
“One mistake leaders make is measuring AI testing tools by volume,” she wrote. “How many tests did it create? How many scripts did it generate? How much time did it save? How many issues did it flag?”
Instead, she suggests a far more demanding benchmark. “A better question is: Did this tool help us find meaningful risk earlier?”
The shift is significant for banking QA teams, many of which are already moving away from measuring automation coverage towards measuring confidence, evidence and risk reduction.
Generating thousands of additional tests offers little value if they fail to expose the software defects that matter most to customers, regulators and the business.
Assurance challenges
Simonova also argued that AI introduces fresh assurance challenges that QA professionals themselves must learn to manage.
“AI does not just casually learn how to perform its various functions, it requires a vast amount of carefully selected data to be fed to it,” she previously said.
“They must ensure that the data being fed to the system is accurate, otherwise the AI model will be trained incorrectly.”
“In software quality, trust cannot be bought through a subscription. It has to be tested.”
– Margarita Simonova
Finally, Simonova also warned that many AI systems remain difficult to interpret. “In general, many people still do not fully understand how AI works. Even AI experts don’t always understand how the system is learning,” she said.
“On top of that, the companies that make AI often try to keep their methods secret. Not fully understanding how AI works can become a big issue.”
Rather than removing humans from testing, Simonova sees AI pushing QA professionals into a more strategic role, reviewing generated artefacts, validating business risks and determining whether automated recommendations genuinely improve release confidence.
“Let AI create the first draft. Let AI summarize information. Let AI suggest risks. Let AI help maintain repetitive checks. Let AI compare screenshots, logs or behavior.”
“But let humans decide what matters,” she stressed.
Simonova’s conclusion is that AI testing tools should ultimately be judged not by how sophisticated they appear, but by whether they improve decision-making.
“The question should not be: ‘Can this tool automate testing?’ The question should be: ‘Can this tool help us make better release decisions?’”
For banks under mounting pressure to deliver software faster without compromising resilience, that may become the defining test of AI-assisted quality engineering.
As Simonova concluded: “In software quality, trust cannot be bought through a subscription. It has to be tested.”
REGISTER TODAY – SIMPLY CLICK HERE
Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
SIGN UP HERE TODAY
REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.