Discovery
epiq
EPIQ × ISA — Discovery Brief

Entity-based discovery,
built around four workstreams.

A short, structured conversation to deepen what we discussed in our first session — so our engineering leadership can shape a tight proposal and a calibrated resource plan before the next working meeting.

01
Email segmentation
Splitting threads & replies into meaningful message units.
02
Recipient & signature extraction
Structured to / from / signature fields as entity signals.
03
Entity normalization
Collapsing name & email variations into canonical form.
04
Entity resolution
Linking representations across datasets into one identity graph.
Takes about 15–20 minutes · Progress auto-saves
For EPIQ · Tech & Product leads
Prepared by ISA Consulting Group
Private & Confidential

01

Who's filling this out

A brief introduction

So we know whose perspective we're hearing — and so we can follow up with the right person if anything needs clarifying.

1.1
1.2

e.g. Tech Lead, Product Lead, VP Engineering, Director of Applied AI.

1.3

02

Project Context

Stakeholders & strategic intent

Confirming the people, the drivers, and the strategic priority behind this initiative.

Stakeholders
2.1
2.2

Tech lead, product lead, eng manager, exec sponsor, etc. Name + role + how they want to be engaged.

2.3

Single name preferred — the person whose "yes" moves us forward.

2.4

e.g. platform, data infra, security, applied AI — we'd love to map the touch points.

Strategic Context
2.5
2.6

Customer demand, competitive pressure, platform consolidation, regulatory, internal mandate — what's the forcing function?

2.7

Helpful to know what worked, what didn't, and what we shouldn't re-litigate.

2.8

03

Problem & Success

Defining the win

What "good" looks like — and how we'll know we got there.

Confirming the problem
3.1

Pretend you're describing it to a new engineer joining the team. We want your framing, not ours.

3.2

What can't users do, or what do they do badly, that this initiative fixes?

3.3

e.g. "find every email between Person A and any external party at Domain X about Topic Y between dates." Concrete examples = sharper backend.

Success metrics
3.4

Adoption, search precision, time-to-answer, cost per matter, customer NPS, engineering productivity — whatever matters most to you.

3.5

What precision floor will the downstream UX accept?

3.6

Internal labelled datasets, prior model accuracy, manual review timings — anything we can anchor to.

3.7

Defensibility, explainability, auditability, latency, cost. Especially relevant for litigation-grade outputs — defensibility often outranks raw precision.

04

Workstream — Segmentation

Splitting emails into message units

Splitting emails containing multiple threads or replies into meaningful message units.

Inputs & scale
4.1
4.2
4.3

Drives parsing complexity. Rough percentages are fine.

4.4

Affects segmentation cues — quoted-reply markers vary by language and client.

What counts as a "message unit"
4.5

Is it always a single email, or sometimes a thread, or sometimes a reply within a quoted chain? The unit definition is the spec.

4.6

Forwarded chains, embedded replies, inline annotations, mixed-language quoting, etc. The pathological 5–10% usually determines production readiness.

4.7

Treated as part of the parent message, separate units, or linked entities?

Current state & quality bar
4.8

Tools, models, vendors, in-house code — what's the baseline we're improving on?

4.9
4.10

If you have a point of view, share it. If you don't, we'll bring one.

05

Workstream — Extraction

Recipients & signatures

Pulling structured to / from / signature fields as high-confidence signals for entity relationships.

Fields & sources of truth
5.1

cc, bcc, return-path, reply-to, signature phone numbers, signature URLs, titles, company names, departments, disclaimers, etc.

5.2

Different reliability profiles. Affects how we weight confidence.

Signature complexity
5.3

Templated corporate, free-form personal, mixed? A few sample sanitized signatures would be 10× more useful than any answer here.

5.4
5.5

Forwarded-chain pollution is a common failure mode — worth flagging if it matters.

Quality, confidence & approach
5.6
5.7
5.8

High-leverage feature for the resolution layer and for defensibility. Most modern pipelines do this.

06

Workstream — Normalization

Canonical form

Collapsing name and email variations, duplicates, and inconsistencies into canonical form.

Scope of normalization
6.1
6.2

More attributes = stronger downstream resolution but more work to QA.

6.3
Canonical form & reference data
6.4

If you have a schema, share it. If not, we'll propose one.

6.5

Corporate directories, customer-supplied custodian lists, public sources.

6.6

Aggressive collapsing reduces dupes but risks false merges. We can tune per use-case.

Quality & ops
6.7
6.8

Human-in-the-loop is often the right answer for litigation-grade workflows.

07

Workstream — Resolution

The identity graph

Linking different representations of the same person/entity across datasets into a unified identity graph.

Linking scope & approach
7.1

Just the email corpus, or also chat, documents, HR data, custodian lists, third-party sources? Cross-corpus linking is often the highest-value use case.

7.2
7.3

e.g. same email address used by two real people over time. The hardest part of identity.

Identity graph
7.4

Native graph DB (Neo4j, Neptune), relational with adjacency tables, search index, or vector store. Trade-offs differ.

7.5

Lookup by ID, fuzzy lookup, traversal, batch enrichment — the API shape drives a lot of internal design.

7.6
7.7

e.g. "this email belonged to Person A from 2018–2022, then Person B." Litigation-grade graphs often need this — a significant design decision if yes.

Quality & governance
7.8
7.9

Often non-negotiable for legal defensibility. Affects model choice (interpretable vs. black-box).

08

Architecture & Stack

How the backend lives

How the backend fits into EPIQ's broader platform — cloud, languages, storage, security, frontend integration.

Platform & runtime
8.1
8.2
8.3
8.4

Internal eDiscovery platform, document review tools, ingestion pipelines, identity / auth, billing, telemetry.

Frontend integration
8.5
8.6
8.7

A Figma share link is ideal. Static images work too — we'll just need a walk-through.

AI / ML stack preferences
8.8

Vendor, on-prem only, no third-party APIs, specific approved models. Common constraint in legaltech given client confidentiality requirements — worth knowing day one.

8.9
Security & compliance
8.10
8.11

Per-matter isolation, regional residency, dedicated vs. shared clusters, customer-managed keys, etc.

09

Team & Engagement

How we work together

Engagement model, sizing, roles, and working model.

Engagement model
9.1

Managed Engineers = ISA-led pod with a tech lead, outcomes-oriented. Staff Aug = engineers embedded into your team day-to-day.

9.2

It's fine to say "not sure" — gut estimate over forced number.

9.3
9.4
9.5
Roles in the pod
9.6

Typical candidates: Tech Lead, Backend Engineer, ML/NLP Engineer, Data Engineer, DevOps/SRE, QA, Solutions Architect, Project Manager.

9.7

e.g. you'll keep product, design, and DevOps in-house — useful so we don't over-staff.

9.8

e.g. signature-extraction experience, identity-resolution at scale, eDiscovery domain depth, specific LLM tooling.

Working model
9.9
9.10
9.11

Standups, demos, planning, on-call expectations. Share what's worked best with past partners.

9.12

ISA's senior team is in Kosovo and Albania (EU hours, strong US overlap). Helpful to know what to lean into and what to avoid.

10

Timeline, Budget & Open Items

Dates, dollars, and anything else

Anything else we should know before designing the proposal.

Timeline
10.1

An approximate quarter helps.

10.2
10.3
10.4
10.5

Customer commitments, regulatory dates, board demos, RFP timing.

Commercials
10.6

Helps us right-size the proposal. Approximate is fine.

10.7
10.8
Risks & open items
10.9

Technical, organizational, commercial, regulatory — whatever's on your mind.

10.10
10.11

11

Final Step

Review & send to ISA

A quick read-through before submitting. You can jump back to any section to edit; your answers are saved.

Thank you — we have it.

Your responses are on their way to ervin.mrishaj@isaconsulting.com and taulant.mehmeti@isaconsulting.com. We'll digest the answers, align internally, and come to the next working session with a draft plan and resourcing recommendation.

If you'd like to share anything else in the meantime, just reply to either of those addresses directly — we appreciate the time you put into this.