The role
0301Design and deploy semantic chunking systems for lengthy, non-uniformly structured legal, tax, and accounting documents
02Build document enrichment pipelines that identify document types, jurisdictions, legal concepts, entities, parties, and other domain-specific metadata
03Develop hierarchical and multi-label document classification systems using both standard and customer-defined taxonomies
04Build LLM-based and traditional NLP information extraction pipelines that identify entities, relationships, citations, references, and key concepts from unstructured content
05Develop knowledge graph construction systems that extract, normalize, connect, and enrich entities, legal concepts, citations, and relationships across large document collections
06Design systems that identify, interpret, and extract insights from complex tabular data embedded within legal, tax, regulatory, and accounting documents
07Create document intelligence capabilities that support downstream search, retrieval, RAG, and agentic AI workflows
08Design robust evaluation frameworks for document understanding systems using expert annotations, synthetic datasets, and production metrics
09Lead technical decisions on document analysis architectures, chunking strategies, extraction methodologies, classification approaches, and knowledge representation frameworks
10Partner closely with engineering teams to deliver scalable, reliable, and production-ready AI systems
11Provide technical leadership and input into AI strategy, platform capabilities, and long-term roadmap decisions
12Mentor applied scientists and machine learning practitioners across the organization