Category Archives: eDiscovery

AI In-Place Classification Takes Enterprise Information Governance to the Next Level

By John Patzakis and Chas Meier

A digital graphic featuring a globe and icons representing various digital communication and data systems, with the title 'AI In-Place Classification Takes Enterprise Information Governance to the Next Level' and a blog attribution for X1.

Ask anyone who has run an enterprise information governance program, and they will tell you the same thing: the defining challenges are scale and massive costs. A regulatory inquiry, privacy audit, data-breach response, M&A data separation, or records-remediation initiative can span mailboxes, file shares, endpoints, Microsoft 365, and other cloud repositories. Enterprise-wide matters routinely involve tens or hundreds of terabytes of unstructured data—and sometimes far more.

With decades of experience in this industry, we have rarely encountered a genuinely small enterprise information governance matter. Once an organization begins operating across its full data estate, it reaches a scale that traditional eDiscovery technology and workflows were never designed to handle.

That reality has shaped—and limited—what organizations could practically do. Conventionally, classifying or analyzing enterprise data required collecting it into a centralized platform and only then performing indexing, classification, search, or review. At multi-terabyte scale, that model becomes problematic on every axis that matters.

Collecting, transferring, and indexing the data can take weeks or months. It creates another large copy of an organization’s most sensitive information outside its original controls. It adds substantial processing, storage, and hosting costs. When the source is a hosted platform such as Microsoft 365, large-scale collection can also encounter service-protection and throttling limits that restrict throughput, often derailing large projects.

As a result, comprehensive classification across distributed enterprise data is most often technically difficult or economically prohibitive. Organizations sampled, narrowed scope, made assumptions, and accepted risks they could not fully see. When a regulator, major incident, or transaction made comprehensive treatment unavoidable, the alternative could be months of work, significant operational disruption, and tens of millions of dollars in expense.

AI In-Place classification changes that operating model. Instead of first moving an entire corpus into a separate AI or review platform, it brings classification to the distributed data through an architecture designed to process content close to its source. The resulting classifications enrich the corresponding indexes without requiring the organization to collect, host, and reindex a second centralized copy of its entire data estate.

That inversion is the difference between applying intelligence only to a small, preselected dataset and making classification practical across the broader enterprise.

X1 Enterprise makes this possible through its distributed micro-index architecture, refined through more than two decades of search and indexing development. Rather than aggregating all enterprise data into a single monolithic index, X1 creates independently managed micro-indexes aligned to meaningful scopes such as user mailboxes, OneDrive accounts, SharePoint sites, endpoints, and file-system locations. Additionally, data can be reviewed in-place for a “spot check” assessment and in-place keyword searches are available as an optional overlay.

X1 has now extended that foundation with patent-pending AI In-Place capabilities. AI models are deployed through the distributed X1 architecture to classify content associated with each micro-index. The classifications are written back as searchable and actionable index attributes. This allows organizations to add AI enrichment without first recollecting the source corpus or constructing an entirely new centralized processing environment.

Because micro-indexes can be created, updated, and enriched independently and in parallel, the architecture distributes work across available infrastructure and limits the effect of any individual update or failure. Classification can also be incorporated into scheduled or incremental index updates, helping organizations maintain a current understanding of their data rather than relying on a one-time snapshot. Millions of AI-enable tags are quickly applied without reindexing.

The practical implications are significant. Organizations can inspect documents, email, messages, and other unstructured content across distributed repositories; identify personally identifiable information, protected health information, payment-card data, privileged material, and other regulated or responsive content; record those findings in the index; and use them to drive search, review, collection, remediation, or policy-based action.

The objective is not to analyze only a sample, a single mailbox, or one repository at a time. It is to apply a consistent classification policy across the relevant enterprise estate.

The economics are equally important. Although the compute and storage requirements are materially reduced, they do not disappear. But the cost curve changes dramatically when classification no longer requires the entire corpus to be transferred, duplicated, centrally hosted, and reindexed before analysis can begin. Existing distributed indexes become the foundation for ongoing enrichment, and only the data that requires further review, collection, or remediation needs to move into downstream systems.

Work that once demanded large, centralized processing environments can instead be distributed across infrastructure designed to absorb enterprise scale. Scale remains a manageable factor, but it no longer has to be the reason an organization cannot perform the analysis at all.

Several governance and compliance use cases that were historically difficult precisely because they are enterprise-scale can therefore become routine. These include:

  • Discovering and remediating PII, PCI, and PHI that has escaped authorized systems and is residing in mailboxes, file shares, endpoints, or collaboration sites where it does not belong.
  • Supporting privacy obligations under GDPR, CCPA, and similar regimes, including data-subject requests, minimization, deletion, and applicable data-location or transfer restrictions.
  • Assessing the scope of a data breach quickly and under regulatory deadlines.
  • Identifying and remediating departed-employee data and insider-risk exposure.
  • Separating data for mergers, acquisitions, and divestitures.
  • Executing records-retention programs and identifying redundant, obsolete, and trivial data across the enterprise.

In each case, the business value has long been clear. The obstacle has been performing the work comprehensively, efficiently, and with minimal additional data movement. AI In-Place classification directly addresses that obstacle.

For years, enterprise information governance has been forced into a largely reactive posture—narrow in scope, expensive, and often a step behind the data. The problem was not a lack of governance objectives; it was a lack of technology capable of supporting them at the scale of modern enterprise information.

AI In-Place classification changes what is operationally and economically practical. It gives organizations a path toward governance that is proactive, comprehensive, and continuously updated across distributed enterprise data—while avoiding the need to centralize another complete copy of the underlying corpus.

There may be no such thing as a small enterprise information governance challenge. There is now a more practical architecture for addressing the large ones.

To learn more about X1 Enterprise and its AI In-Place capabilities, visit x1.com or contact sales@x1.com.

Leave a comment

Filed under AI In-Place, Best Practices, Cloud Data, Data Governance, ECA, eDiscovery, eDiscovery & Compliance, Enterprise AI, Enterprise eDiscovery, ESI, Information Governance, Information Management, m365

De-NISTing in eDiscovery: A Costly Provision That Shouldn’t Be in Model Orders in the First Place

By John Patzakis

A model eDiscovery order I recently came across from a federal district court issued by a respected judge included a provision requiring parties to de-NIST their files in the course of eDiscovery production. On its face, this may seem like a reasonable technical requirement to some practitioners. But this provision reflects a fundamental misunderstanding of how proportional, targeted eDiscovery collection should work — and it points to a broader problem in our industry that deserves some attention.

For those unfamiliar with the term, de-NISTing refers to the process of filtering out known, irrelevant system files from a forensic collection using the National Institute of Standards and Technology’s reference database of known file signatures. The NIST database catalogs hundreds of thousands of known operating system files, executables, DLL files, and other system-generated data that have no evidentiary value whatsoever. De-NISTing removes these files from a collection so that reviewers are not burdened with wading through mountains of irrelevant system data. The reason you need to de-NIST in the first place is because you collected a full-disk image — capturing everything on the drive, relevant or not.

And that is precisely the problem with requiring de-NISTing in a model eDiscovery order. As I have written extensively, including in our recent white paper on proportionality in eDiscovery, courts have consistently held that full-disk imaging is not the appropriate default for civil litigation collections. Going all the way back to Deipenhorst v. City of Battle Creek in 2006, courts have warned that imaging a hard drive results in the production of massive amounts of irrelevant — and potentially privileged — information. More recently, in Motorola Solutions v. Hytera Communications Corp., the court emphasized that forensic examination of a party’s computers “is no routine matter” and that courts must use caution to avoid unduly impinging on privacy interests. A model order that presupposes full-disk imaging by requiring de-NISTing is, at minimum, inconsistent with this well-established body of case law.

The 2015 amendments to Federal Rule of Civil Procedure 26(b)(1) established a clear six-pronged proportionality framework for eDiscovery, requiring parties and courts to weigh factors including the importance of the issues at stake, the amount in controversy, the parties’ resources, and whether the burden or expense of proposed discovery outweighs its likely benefits. Courts have taken these amendments seriously and have consistently limited overbroad discovery requests on proportionality grounds. A blanket model order requirement to de-NIST implicitly endorses a collect-everything methodology that runs counter to the proportionality principles embedded in Rule 26(b)(1) and the extensive case law that has developed around it.

So how does a provision like this end up in a model court order? The answer, I believe, lies in the undue influence that certain eDiscovery service providers have had on collection practices and, ultimately, on the drafting of court orders and guidelines. Some service providers have a clear financial incentive to collect as much data as possible, since their fees are calculated on a per-gigabyte basis — meaning the more data collected, processed, and hosted, the higher the bill. This volume-based business model has shaped industry “best practices” in ways that favor over-collection, and that mindset has quietly seeped into the thinking of some federal judges and the model orders they issue. What gets dressed up as technical diligence is, in many cases, simply an artifact of a business model that profits from excess.

If you are conducting a properly scoped, targeted eDiscovery collection that is consistent with the principles of proportionality — as the Federal Rules and overwhelming case law require — there is simply no reason to de-NIST. A targeted collection does not reach system files, executables, DLLs, or other non-user-generated data in the first place. You are collecting potentially relevant ESI from identified custodians, scoped by search terms, date ranges, file types, and data sources. You never touch the data that de-NISTing is designed to filter out, which means the entire de-NISTing step — and its associated cost and processing time — is unnecessary overhead born entirely of an overbroad collection methodology.

This is precisely the approach built into X1 Enterprise, which enables legal and IT teams to conduct targeted, remote collections across large numbers of custodians without ever capturing the system-level data that necessitates de-NISTing. X1 Enterprise collects only the user-generated, potentially relevant ESI within defined parameters, preserving full metadata integrity and maintaining a documented chain of custody — satisfying every requirement for forensic soundness without the bloat, expense, and proportionality concerns of full-disk imaging. In an era where courts are increasingly scrutinizing eDiscovery costs and demanding proportionality, practitioners and judges alike should be asking not how to manage the mess created by over-collection, but how to avoid creating that mess in the first place.

Leave a comment

Filed under Best Practices, Case Law, Cloud Data, Cybersecurity, Data Audit, Data Governance, eDiscovery, eDiscovery & Compliance, Enterprise eDiscovery, GDPR, Information Governance, Information Management

AI Without Data Movement: X1’s Webinar Reveals the Future of Secure Enterprise AI

By John Patzakis

X1’s recent webinar announcing the availability of true “AI in-place” for the enterprise was both highly attended and strongly validated by the audience response. The session did more than introduce a new feature; it articulated a fundamentally different architectural approach to enterprise AI—one designed explicitly for security, compliance, and scalability in complex, distributed environments. Our central message was simple: enterprise AI adoption has been constrained not by lack of interest, but by architectural and security requirements that existing platforms have failed to address.

That reality was most powerfully captured in a quote shared on the opening slide from a Fortune 100 Chief Information Security Officer, which set the tone for the entire discussion:

“Normally AI for infosec and compliance use cases is a non-starter for security reasons, but your workflow and architecture is completely different. This allows us – all behind our firewall — to develop our own models that are trained on our own data and customized to our specific security and compliance use cases and deployed in-place across our enterprise.”

This endorsement crystallized the webinar’s core insight: AI becomes viable for the most sensitive enterprise use cases only when it is deployed where the data already lives, rather than forcing data into external or centralized systems.

The technical foundation that makes this possible is X1’s micro-indexing architecture. Unlike traditional platforms built on centralized, resource-intensive indexing technologies, X1 deploys lightweight, distributed micro-indexes directly at the data source. This allows enterprises to index, search, and now apply AI analysis without mass data movement. As emphasized during the webinar, centralized indexing is not just expensive and slow—it is fundamentally misaligned with how modern enterprise data is distributed across file systems, endpoints, cloud platforms, and collaboration tools.

The session then highlighted how this architectural distinction resolves a long-standing problem in discovery, compliance, and security workflows. Legacy platforms require organizations to collect and centralize data before they can analyze it, introducing delays, high costs, and significant risk exposure. X1 reverses that workflow. By enabling visibility and AI-driven classification before collection, organizations can make informed, targeted decisions—collecting only what is necessary, remediating issues in-place, and dramatically reducing both risk and operational overhead.

The discussion also demystified large language models (LLMs), explaining that while model training is compute-intensive, models themselves are increasingly commoditized and portable. Critically, LLMs require extracted text and metadata— processed from native files—to function. This aligns perfectly with X1’s existing capability, as text and metadata extraction are already integral to our micro-indexing process. AI models can therefore be deployed alongside these indexes, operating in parallel across thousands of data sources with massive scalability.

The conversation then connected this architecture to concrete, high-value use cases. In eDiscovery, AI in-place enables faster early case assessment and proportionality by analyzing data where it resides. In incident response and breach investigations, security teams can immediately scope exposure across distributed systems without waiting months for data exports. For compliance and governance, AI models can continuously identify sensitive data, enforce retention policies, and surface risk conditions that were previously impractical to monitor at scale.

In addition to a live product demo showcasing this new capability, we concluded the webinar with several clarifying points and announcements. First, we emphasized that X1 does not access, monetize, or host customer data. Also, AI in-place is not an experimental add-on but an enhancement to a proven, production-grade platform. And notably, there is no additional licensing cost for the AI capability itself—customers simply deploy models within their own environment. With proof-of-concept testing beginning shortly and production deployments targeted for April 2026, the webinar made clear that AI in-place is not a future vision, but an imminent reality for the enterprise.

You can access a recording of the webinar here, and to learn more about X1 Enterprise, please visit us at X1.com.

Leave a comment

Filed under Best Practices, Corporations, Cybersecurity, Data Audit, Data Governance, ECA, eDiscovery, eDiscovery & Compliance, Enterprise AI, Enterprise eDiscovery, ESI, Information Governance

Why Most SaaS Architectures Fall Short for Enterprise-Grade AI

By John Patzakis and Chas Meier

SaaS Architectures Fall Short for Enterprise-Grade AI

As organizations accelerate adoption of AI to support legal, compliance, security, and business operations, one principle is becoming clear: the underlying deployment architecture matters as much as the model itself. Many enterprise AI initiatives fail not because the technology is immature, but because the environment in which it operates was never designed for high-volume, sensitive, or tightly regulated use cases.

Traditional multi-tenant SaaS architectures—where numerous customers share the same provider-controlled environment—excel at delivering standardized, lower-risk business applications. But applying that same model to AI workloads involving privileged, regulated, or company sensitive data introduces material limitations in governance, security, performance, and operational feasibility.

Below are the core architectural constraints that legal, IT, and security leaders consistently raise as they evaluate AI strategies.

  1. Data Governance, Privacy, and Regulatory Control
    Most commercial SaaS AI platforms require customer data—or derivative artifacts such as embeddings, logs, or temporary working sets—to be processed within the provider’s environment. Even with strong encryption and contractual controls, this shift of data outside the enterprise’s controlled boundary introduces challenges that many legal and security teams cannot accept.

    Key concerns include:
    Loss of direct data sovereignty. Once data is inside a vendor’s multi-tenant environment, the organization no longer controls how it is stored, moved, or isolated.
    Jurisdiction and residency risks. Multi-tenant SaaS services often replicate or route data across regions for load or resilience purposes, complicating GDPR, HIPAA, ITAR, or sector-specific compliance requirements.
    Governance of secondary artifacts. AI systems often generate embeddings, caches, metadata, and diagnostic logs. Ensuring these artifacts adhere to the same retention, destruction, and legal hold rules become significantly more complex in a shared environment.

    For legal departments, eDiscovery teams, and CISOs, these factors create an expanded compliance burden that is often disproportionate to the value of outsourcing AI workloads.
  2. Assurance of Isolation and Auditability
    Large enterprises increasingly demand verifiable guarantees—not merely assurances—that:
    • Their data is isolated from other tenants
    • Their information is not used for model training unless explicitly authorized
    • Every transaction is auditable and traceable
    • No shared services introduce inadvertent cross-tenant visibility

    While reputable AI providers enforce strong separation controls, multi-tenant architecture inherently increases the assurance burden. The organization must rely on the vendor’s internal controls, certifications, and change management practices—none of which it can independently verify.

    For regulated entities, this can be an unacceptable dependency, particularly where privileged legal data, sensitive communications, or proprietary research is involved.
  3. Performance and Scalability Under AI Workloads
    AI inference and large-scale analysis require sustainable compute performance. Multi-tenant environments, by design, pool capacity across customers. Even when quotas or isolation tiers exist, resource contention and dynamic scaling can introduce variability.

    For enterprise workloads—such as legal investigations, regulatory responses, internal audits, or global compliance monitoring—performance variability translates directly into operational delays and risk.

    Organizations routinely raise:
    Deterministic performance requirements for time-sensitive matters
    Workload isolation needs when running tens of thousands of queries or document classifications
    The high cost of dedicated capacity tiers in third-party SaaS models

    These are structural limitations, not configuration issues.
  4. Data Movement, Transfer Overhead, and Operational Disruption
    Before any SaaS-based AI workflow begins, enterprises must stage or transfer large volumes of data—including emails, documents, chat messages, or historical repositories—into the vendor’s cloud environment.

    This poses several obstacles:
    Time and bandwidth constraints when transferring terabytes or petabytes
    Chain-of-custody and legal hold considerations during data movement
    Jurisdictional restrictions when data cannot transit or be stored outside specific regions
    Ongoing synchronization challenges as new data is generated

    For legal, compliance, and security teams, these issues often make multi-tenant SaaS unsuitable for high-value unstructured data.
  5. Limited Customization and Restricted Model Control
    Most multi-tenant AI SaaS offerings operate within a shared, standardized stack. This limits an enterprise’s ability to:
    • Tailor models to domain-specific content or workflows
    • Implement custom inference pipelines
    • Integrate internal security, monitoring, or policy engines
    • Maintain visibility into how models process and route sensitive information

    For departments handling privileged, confidential, or regulated data, this lack of deep configurability hampers both innovation and risk mitigation.

The Industry Shift Toward AI-in-Place Architectures
To address these concerns, organizations are increasingly adopting AI-in-Place models—deploying AI capabilities directly onto systems, repositories, and environments they already control.

AI-in-place allows enterprises to:
• Keep all source data behind the firewall or within their private cloud tenancy
• Maintain full sovereignty over models, embeddings, logs, and derived artifacts
• Enforce internal security, retention, and access policies without exception
• Optimize performance around their own infrastructure and workflows
• Reduce compliance complexity by avoiding data egress entirely

This architectural shift reflects a maturing understanding: the value of AI is maximized only when it can operate where sensitive data already resides.

X1 Enterprise: A Modern Foundation for AI-in-Place
X1 Enterprise—with its patented distributed micro-indexing architecture—has emerged as a leading platform for organizations adopting AI-in-Place strategies.

X1 enables:
In-place analysis without data movement
Deploy LLMs, embeddings, and AI pipelines directly to endpoints, repositories, and cloud data sources—without exporting or copying sensitive content.
Enterprise-wide visibility across unstructured data
Email, documents, chat, archives, and cloud sources can be searched, tagged, classified, and analyzed at scale from a single federated index.
High-assurance governance
All data remains within the enterprise’s security boundary or isolated single-tenant cloud, supporting legal holds, audits, discovery, and regulatory requirements.
Scalable performance tailored to the enterprise’s environment
Micro-indexing distributes compute to where data lives, eliminating bottlenecks inherent in centralized SaaS architectures.

For legal, IT, and security leaders seeking to implement AI responsibly, X1 provides a practical and compliant path forward.

See AI-in-Place in Action
We invite you to join our upcoming webinar on Wednesday, December 10, where our team will present:
• A detailed look at X1’s new AI-in-Place capabilities
• Architectural considerations for legal, IT, and CISO stakeholders
• A live demonstration of enterprise-scale AI applied directly to live data sources

Register here to secure your spot.

Leave a comment

Filed under Best Practices, Cloud Data, Corporations, Cybersecurity, Data Audit, Data Governance, eDiscovery, eDiscovery & Compliance, Enterprise AI, Enterprise eDiscovery, Information Governance, SaaS

X1 Brings “AI In-Place” to the Enterprise—A Major Breakthrough for Secure, Scalable AI Deployment

By John Patzakis

Our latest announcement represents a true inflection point in enterprise AI. With X1 Enterprise’s newly introduced capability for AI in-place, organizations and their service providers will, for the first time, be able to deploy and execute large language models (LLMs) directly where enterprise data lives—without moving or copying that data.

This is more than a product enhancement; it is a fundamental shift in how AI is applied across the enterprise.

The Foundation: Efficient Text Extraction Is Critical for AI
Large language models (LLMs) are the core engines that power today’s AI revolution. These models rely entirely on textual input to perform reasoning, summarization, search, and analysis. That is why text extraction is the critical first step. LLMs can only operate once another process extracts the text from emails, documents and chats. Traditionally, that meant copying or exporting data to external systems hosted by third party vendors, a process fraught with risk, cost, and compliance challenges.

Solving the “Data Movement Problem” for Enterprise AI
So, the key barrier to enterprise AI adoption has been the reluctance to move sensitive corporate data to external AI platforms. Whether for security, governance or cost reasons, most enterprises simply cannot send their data outside their environment.

X1’s innovation solves that problem head-on. Instead of shipping sensitive data out to an AI system, X1 brings the AI to the data. Enterprises can now deploy their own proprietary models or open-source LLMs within the secure perimeter of their existing infrastructure, whether on premises or in the cloud. X1’s index-in-place architecture performs the text extraction and indexing where the data resides. By extending that same principle to AI—forward-deploying LLMs directly to enterprise data sources—X1 now enables AI in-place. The result: organizations can apply the analytical power of LLMs across their data without ever moving it.

Once the LLMs are deployed into the X1 micro-indexes, X1 will then auto-apply AI-informed tags, which a user can query globally from a central console and act upon through targeted data collection or remediation. Imagine petabytes of data on file servers, laptops M365 and other sources all AI-classified and then queried and collected on a highly targeted basis.

This means enterprises can now unlock powerful new use cases no matter the scale—AI-assisted compliance, risk monitoring, GRC audits, eDiscovery, and more—while maintaining full control of their data and eliminating the need for costly, risky data transfers.

Enabling Collaboration Between Enterprises and Their Advisors
William Belt, Managing Director and Consulting Practice Leader at Complete Discovery Source, described the impact succinctly:

“Enabling AI in-place where our corporate client’s data lives is game-changing. We look forward to working with our clients to deploy AI models that are either pre-trained or customized for a specific matter or compliance requirement utilizing the X1 Enterprise platform.”

This capability creates a new bridge between corporations and their professional advisors—consulting firms, law firms, and service providers—who can now collaborate directly with their clients to develop, fine-tune, and deploy customized AI models for specific business or legal needs.

Rather than relying on generic cloud-based AI tools, organizations can now build targeted, matter-specific LLMs that are tuned to their unique data and compliance requirements, all executed securely in-place through the X1 Enterprise Platform.

A New Era for Enterprise AI
With this release, X1 is redefining the architecture of enterprise AI. Its ability to perform distributed micro-indexing and in-place AI analysis across global data sources enables secure, scalable, and cost-effective intelligence—without ever duplicating or relocating sensitive data.

For enterprises and their partners, this represents a new era of possibility: true AI at enterprise scale, in-place.

X1 will host a webinar on Wednesday, December 10, featuring a detailed overview of this new capability and a live demonstration. You can register here.

Leave a comment

Filed under Cloud Data, Corporations, Cybersecurity, eDiscovery, eDiscovery & Compliance, Enterprise AI, Enterprise eDiscovery, Information Governance, m365