Annotaurus
PTI text annotation platform for assigning tasks, collecting labels, and connecting human review to downstream extraction workflows.
Portfolio
Public notes on selected research, teaching, software, and language technology projects.
PTI text annotation platform for assigning tasks, collecting labels, and connecting human review to downstream extraction workflows.
Course operations tooling for managing teaching workflows, student-facing materials, protected instructor resources, and authenticated course infrastructure.
Mobile-friendly book intake, metadata extraction, valuation evidence, and buyer sharing for inherited, rare, signed, and resale-relevant books.
Android-first DC-1 app for reading PDFs and EPUBs, handwritten annotation, free-form notes, and capture into Exocortex.
TLM event discovery and ingestion infrastructure for extracting, normalizing, and curating live-music records from hundreds of disparate venue and promoter websites.
Research on executable, editable extraction programs that combine linguistic structure, neural assistance, provenance, and human review.
Machine reading, information extraction, knowledge graphs, and grant-supported applied NLP systems.
See allResearch on executable, editable extraction programs that combine linguistic structure, neural assistance, provenance, and human review.
Neural and LLM-assisted program synthesis systems for generating explainable, editable extraction rules from examples.
Applied language technology for making trusted public health information easier to search, summarize, translate, and reuse.
Research on modeling stylistic evidence for attribution, verification, controlled rewriting, and language variation.
Research on tokenization and representation choices for language communities poorly served by mainstream pretrained vocabularies.
Student collaboration on pronunciation practice and feedback support for learners of Arabic.
Research on how social-group framing and subtle linguistic choices shape descriptions of people, actions, responsibility, and social meaning.
Biomedical and scientific machine-reading work around extracting events, assembling evidence, and building inspectable knowledge structures.
Specification-driven tools, publishing systems, teaching infrastructure, and personal knowledge work software.
See allCourse operations tooling for managing teaching workflows, student-facing materials, protected instructor resources, and authenticated course infrastructure.
Android-first DC-1 app for reading PDFs and EPUBs, handwritten annotation, free-form notes, and capture into Exocortex.
Mobile-friendly book intake, metadata extraction, valuation evidence, and buyer sharing for inherited, rare, signed, and resale-relevant books.
Automation infrastructure for coordinating repository work, issue context, CI feedback, and development operations.
PTI text annotation platform for assigning tasks, collecting labels, and connecting human review to downstream extraction workflows.
TLM event discovery and ingestion infrastructure for extracting, normalizing, and curating live-music records from hundreds of disparate venue and promoter websites.
Advancing Indigenous Language Technologies work around community-governed language technology, curation, search, and access-aware infrastructure.
See allCommunity-facing dictionary, search, and curation infrastructure for Indigenous language learning and language-program needs.
Startup and client-facing applied NLP work where public descriptions emphasize safe, high-level system roles.
See allLLM-driven structured information extraction over transcribed interviews, using private model deployment and an explicit target schema.
Vision-based PI&D understanding for engineering PDFs, combining document analysis, extraction, and graph-oriented structure.
MCP surface for deep DOCX editing against document styles, templates, and structured editing constraints.
Literature-based discovery platform joining multi-domain causal extractions into a searchable and editable knowledge graph.
Neural and LLM-assisted program synthesis systems for generating explainable, editable extraction rules from examples.