25 April 2026 · 9 min read · The Implementation Layer

We scanned 5 AI frameworks for EU AI Act compliance. Here’s what 389 patterns found.

Scan performed with Regula v1.7.0 (389 patterns at the time; current version has 423).

Guides

Your compliance surface doesn’t start with the code you write. It starts with the framework you import. We ran Regula against the source code of PyTorch, HuggingFace Transformers, LangChain, LlamaIndex, and CrewAI — five frameworks that collectively power most of what gets called “AI” in production today.

Why scan frameworks, not apps?

Framework code can expose technical patterns that an application team should review. Importing a library does not by itself transfer an Article 14 or Article 15 obligation to every application; legal relevance depends on the resulting system, its classification, use, and implementation.

Our previous post scanned 10 open-source AI applications and found 553 findings. That told us what developers are building. This post asks a different question: what does the foundation layer look like?

Frameworks don’t make deployment decisions. They don’t decide whether a chatbot runs in a hospital or a toy shop. But they establish patterns that propagate into every application built on top of them. A model-loading library that uses pickle deserialization carries a cybersecurity surface (Article 15) into every project that imports it. An agent framework that pipes LLM output to subprocess.run carries a human oversight surface (Article 14) into every agent built with it.

So we scanned them.

The frameworks

We scanned PyTorch, HuggingFace Transformers, LangChain, LlamaIndex, and CrewAI — five frameworks with a combined 492,717 GitHub stars and 19,426 source files. Together they power most of what ships as AI in production today, from model training and serving to agent orchestration and retrieval-augmented generation.

Framework Stars What it does Source files Findings
HuggingFace Transformers 159,867 Pre-trained model library (NLP, vision, audio) 4,255 175
LlamaIndex 48,882 Data framework for LLM applications 4,594 163
PyTorch 99,411 Deep learning framework 7,010 93
CrewAI 49,778 Multi-agent orchestration framework 1,104 78
LangChain 134,779 LLM application framework 2,463 53
Total 492,717 19,426 562
Guides

Scanned with Regula v1.7.0 (389 risk patterns across EU AI Act, OWASP LLM Top 10, and OWASP Agentic Security). Under 4 seconds per framework except PyTorch (7,010 files, ~30 minutes).

What we found

562 findings across 19,426 source files. AI security accounted for 54.1%, followed by agent autonomy (18.5%), high risk (14.2%), limited risk (10.1%), and credential exposure (3.0%). No prohibited-practice patterns were found in this scan. Checked-in scan results.

Category Findings % What it means
AI security 304 54.1% Unsafe deserialization, sensitive data in prompts, unbounded generation — maps to OWASP LLM Top 10 and Article 15
Agent autonomy 104 18.5% LLM output flowing to system commands, file writes, HTTP calls without human oversight gates
Limited risk 57 10.1% Chatbot interfaces, synthetic content generation — Article 50 transparency obligations
High risk 80 14.2% Patterns associated with employment, biometric, education, or safety-critical contexts
Credential exposure 17 3.0% API keys or tokens detected in source — Article 15 cybersecurity relevance
Prohibited 0 0% No prohibited practice patterns found in any framework

Source data for this table.

Compare this to the application-level scan: when we scanned 10 AI apps, agent autonomy was the top category at 56.6%. In frameworks, AI security dominates at 54.1%. That makes sense. Frameworks handle the plumbing (model loading, data serialization, API wiring). Applications make the deployment decisions.

Zero prohibited findings. Frameworks don’t implement social scoring or subliminal manipulation. If you’re worried about Article 5, look at what you’re building with the framework, not the framework itself.

Each framework carries a different compliance profile

Within this scan's own finding taxonomy, each framework concentrated matches differently. Those percentages describe scanner findings, not the share of legal obligations or compliance risk. Application context and Article 6 classification remain untested.

We didn’t expect the differences to be this stark.

HuggingFace Transformers: cybersecurity surface

113 of 175 findings are AI security patterns. Transformers loads, saves, and converts models across dozens of formats. That means deserialization code paths and model weight handling, which trigger Article 15’s cybersecurity requirements if the system qualifies as high-risk.

It also had 28 high-risk findings, the highest count of any framework. Transformers is used in computer vision, speech recognition, and NLP classification, all of which can land in Annex III high-risk categories depending on how they’re deployed.

If you are building on Transformers, the scan suggests reviewing model integrity, deserialisation, and provenance. It does not establish that cybersecurity is the project's main legal concern.

CrewAI: human oversight surface

56 of 78 findings are agent autonomy patterns. CrewAI’s entire purpose is to let AI agents take actions, and the scan reflects that: LLM output flowing to tool execution, task delegation between agents, autonomous decision loops.

If you deploy a CrewAI-based system in a high-risk context, Article 14 requires that humans can effectively oversee it. Confirmation gates, audit trails, the ability to interrupt agent decisions.

CrewAI provides some of these primitives (human input steps, callbacks). The scan can’t tell whether your specific pipeline actually uses them. That’s on you.

PyTorch: deployment-context and cybersecurity surface

93 findings, split between AI security (53) and high-risk (38). PyTorch’s 38 high-risk findings are the highest of any framework, which makes sense: it’s the infrastructure behind computer vision, speech processing, and classification systems that regularly land in Annex III high-risk categories.

The AI security findings come from model serialization (torch.save/torch.load using pickle), CUDA memory handling, and distributed training patterns. The high-risk findings come from PyTorch’s use in biometric processing, safety-critical systems, and education assessment contexts.

Notably, only 2 agent autonomy findings. PyTorch doesn’t orchestrate agents; it trains and serves models. The compliance surface is cybersecurity and deployment context, not oversight.

LlamaIndex: mixed surface with credential exposure

163 findings, the highest count. 98 are AI security, mostly from its data connectors and document processing pipelines. LlamaIndex connects to dozens of external data sources (databases, APIs, file systems, cloud storage), and each connection point is an attack surface.

The 16 credential exposure findings are the interesting ones. LlamaIndex’s integration layer handles API keys for vector databases, LLM providers, and data sources. Some of these appear in example code and test fixtures. Developers copy-paste examples. If the example has something that looks like a real API key, it ends up in production.

28 agent autonomy findings come from its query engine and agent capabilities, which route user queries to retrieval and generation pipelines.

LangChain: relatively clean for its size

The recorded scan found 53 findings across 2,463 source files. LangChain’s monorepo splits functionality into small packages, so patterns do not concentrate in one place. Checked-in scan results.

Top indicators: sensitive information disclosure (18) and system command execution (9). The agent patterns exist (LangChain has tool-calling and agent execution) but they’re spread across packages rather than concentrated in a single orchestration layer.

What this means for teams choosing frameworks

Framework choice affects technical architecture and therefore the evidence worth reviewing. It does not automatically create a specific EU AI Act obligation. Confirm system scope, role, classification, and actual safeguards before drawing a legal conclusion.

The framework-level story is AI security, not agent autonomy. When we scanned applications, agent autonomy dominated. When we scanned frameworks, AI security dominated. Frameworks create the security surface. Applications create the oversight surface. Different conversations, different teams.

Article 5 is an application-layer concern. No framework ships social scoring or subliminal manipulation. If you’re evaluating prohibited-practice risk, you’re evaluating what you build, not what you build with.

One thing we didn’t expect: credential exposure in example code. LlamaIndex’s 16 credential findings are mostly in examples and templates. Developers copy-paste examples. Framework maintainers should treat example code with the same security discipline as library code.

Why 389 patterns matter

Some scanning tools run 39 checks. Regula now runs 423 tiered risk patterns (389 at the time of this scan), covering the EU AI Act (Articles 5, 9-15, 50, 51-55), OWASP LLM Top 10, OWASP Agentic Security, and 18 credential patterns. The difference shows up most in the AI security category, where deserialization risks and prompt injection surfaces need specific regex matchers to avoid false positives.

At 39 checks, a model-loading framework looks clean. At 389, you can see the cybersecurity surface that Article 15 actually cares about.

Run this yourself

Install Regula with pip, clone a framework, and scan it in one command. Regula is open-source, has zero core dependencies, and its local scan engine does not upload project files. Full JSON scan output for every framework in this post is published in the repository for independent verification.

pip install git+https://github.com/kuzivaai/getregula.git@main
git clone --depth 1 https://github.com/langchain-ai/langchain.git
regula check langchain

Regula is open-source, has zero core dependencies, and processes scan inputs locally without uploading project files. The full scan JSON for each framework in this post is available in the repository for independent verification.

Last reviewed: 14 August 2026 · Author: Regula maintainers · Not legal advice. Regula identifies risk indicators for developer review.

Scans were run using Regula v1.7.0 against shallow clones of each framework’s default branch on 25 April 2026. Star counts are as of the same date. Source file counts include Python, JavaScript, TypeScript, C/C++, and Jupyter notebooks. Findings reflect code-level patterns, not deployed system risk classifications — the EU AI Act regulates systems as deployed, not code in isolation.

Regulatory update, 4 August 2026: Regulation (EU) 2026/1744 is enacted and in force. It sets 2 December 2027 for the Annex III path and 2 August 2028 for the Annex I product-embedded path. This April 2026 scan remains a historical benchmark; findings are code indicators, not legal classifications.

Related reading

Discuss on Hacker News  ·  Guides View on GitHub