PentesterFlow Automates Pentesting With Human-in-the-Loop AI

Dominic Jainy brings a wealth of experience in artificial intelligence and machine learning to the complex field of offensive security. As a professional who has navigated the integration of emerging technologies across various sectors, he possesses a keen eye for how agentic systems can either empower or endanger a security posture. Today, we sit down with Dominic to discuss PentesterFlow, a new open-source tool that aims to redefine the penetration testing lifecycle by balancing autonomous action with human oversight.

This discussion covers the evolution of AI-driven security workflows, emphasizing the shift toward human-in-the-loop architectures that mitigate the risks of hallucination and unauthorized command execution. Dominic breaks down the mechanics of local continuous learning, the seamless bridge between the command line and professional tools like Burp Suite, and the necessity of evidence-backed reporting in modern bug hunting. We also explore the practicalities of setting up local large language models and the importance of maintaining strict analyst control during sensitive engagements.

Many automated tools struggle with context retention and hallucinated findings. How does a local continuous learning system change the way an AI agent handles a complex pentest?

PentesterFlow solves the persistent frustration of an AI “forgetting” the layout of a target after a few prompts by utilizing project-specific and personal intelligence files. These files act as a long-term memory bank that silently archives successful workflows, coverage gaps, and lessons learned from failed assumptions without requiring the user to retrain a massive underlying model. It feels remarkably smooth to see the tool recall a specific API quirk from an hour ago and pivot its strategy accordingly rather than repeating the same mistake. By automatically redacting secrets and deduplicating data before it hits the disk, the system maintains a high-signal environment for the LLM. This local intelligence ensures that whether you are using a local Ollama instance or a hosted Gemini API, the agent gets smarter with every HTTP request it sends to the target.

Trust is a major barrier when it comes to autonomous security tools. How does the human-in-the-loop design specifically address the fears of running AI against production systems?

The human-in-the-loop design is the emotional anchor for any security engineer who has ever worried about an automated script knocking a critical database offline. PentesterFlow enforces a permission-gated execution model where every sensitive shell command or invasive probe requires an explicit sign-off from the analyst before it runs. You can almost feel the relief of knowing there is a “stop” button and a clear preview before a potentially catastrophic pattern is executed against a scoped target. While there is a “YOLO mode” available for those working in isolated lab environments who want to let the AI run wild, the default setting ensures that credential redaction and SHA-256 checksum verification keep the operation professional and safe. It effectively transforms the AI from an unpredictable loose cannon into a disciplined digital apprentice that respects the strict boundaries of the engagement.

For a bug hunter, the transition from discovery to documentation is often the most tedious part. How does this tool streamline the workflow between finding a vulnerability and generating a report?

One of the most tangible benefits of this tool is the Burp Suite bridge, which allows a tester to move effortlessly between manual interception and AI-assisted enumeration. Instead of juggling disparate windows and manually copying data, you can send captured traffic directly into the CLI and then import confirmed findings back into Burp as actionable issues. The reporting side is equally robust, automatically generating Markdown files that include copy-pasteable curl commands and clear, evidence-backed proofs of concept. This evidence-based approach removes the guesswork, providing the exact impact and remediation steps needed for a professional-grade report. It is a deeply satisfying experience to watch a high-severity IDOR vulnerability get validated and documented in real-time with almost zero manual data entry required.

The tool mentions a wide range of “skills” like SSRF, SSTI, and GraphQL. How does the agent actually apply these during the reconnaissance and enumeration phases?

The breadth of built-in skills is impressive, covering everything from Supabase exploits and race conditions to subdomain takeover and deserialization. Once the analyst sets a target URL with a simple command, the agent can be instructed in plain English to “test the orders API for broken access control” or similar objectives. It then loads the relevant “webvuln” skill, plans a sequence of actions, and executes real offensive-security tools to probe the target. You can observe the agent sending targeted HTTP requests and automatically confirming vulnerabilities with technical evidence, which is far more reliable than a simple scanner. Watching the terminal light up with a confirmed finding after a series of automated but supervised steps is a glimpse into how much faster the entire lifecycle can become.

What is your forecast for the role of agentic AI in offensive security over the next few years?

I believe we are entering an era where the “lone wolf” pentester will be replaced by a conductor leading an orchestra of specialized AI agents. In the next few years, the focus will shift away from basic automation toward these sophisticated, evidence-backed systems that can handle the heavy lifting of reconnaissance and repetitive enumeration. We will likely see a drastic reduction in the time spent on “grunt work,” allowing human experts to focus purely on the creative, high-level logic flaws that AI still struggles to grasp. However, the tools that ultimately succeed will be the ones that prioritize transparency and human approval over total autonomy, because trust remains the most valuable currency in the cybersecurity industry.

Explore more

Can the Poco M8 Power Last Three Days on a Single Charge?

The relentless evolution of mobile hardware has reached a critical juncture where the primary concern for modern consumers is no longer pure processing speed but the longevity of a single charge under demanding conditions. As the industry moves into the second half of 2026, manufacturers are increasingly pivoting toward power management solutions that promise to untether users from their wall

Trend Analysis: UK Workplace Harassment Regulations

The corporate landscape in the United Kingdom is currently undergoing a transformative shift as the legal threshold for preventing workplace harassment moves from a reactive posture to a stringent proactive mandate. This legislative evolution forces organizations to move beyond check-the-box compliance and take ownership of employee safety. The transition from the “reasonable steps” standard to a rigorous “all reasonable steps”

Visa and LianLian Global Pilot First AI Agent B2B Payment

Introduction The financial landscape in Greater China recently witnessed a monumental transformation as autonomous intelligence moved beyond mere administrative assistance to handle complex business transactions independently. This pilot program represents the first successful live B2B transaction conducted via an agentic system in the region. By utilizing the LoopXPay agent, Visa and LianLian Global demonstrated that artificial intelligence can navigate the

How Does Foxit PDF Reader Allow Full SYSTEM Takeover?

Introduction Cybersecurity professionals frequently observe that the most profound dangers often stem from a misplaced trust in legitimate software applications that possess deep operational access to the underlying operating system. The discovery of the vulnerability identified as CVE-2026-57239 serves as a stark reminder that even widely utilized productivity tools like Foxit PDF Reader can inadvertently become conduits for a complete

Oracle Advances Hybrid Cloud and Private AI Integration

The historical momentum that once pushed nearly every enterprise database toward the centralized public cloud has finally encountered the immovable reality of data sovereignty and processing latency. For years, the digital transformation narrative focused exclusively on moving assets to massive hyperscale data centers. However, the rise of sophisticated artificial intelligence and strict global privacy laws has forced a rethink of