AI summary: Designs and owns the core AI system architecture, agentic workflows, and retrieval pipelines that power multi-product AI agents in production.
Principal AI Engineer
BuzzBoard · Remote (WFH) · Engineering / AI
The role
BuzzBoard builds AI products for the B2SMB market — helping agencies, media companies, and sellers understand small businesses and market to them at a level of personalization that wasn’t previously economical. Our platform is built on frontier models end to end. Not AI features bolted onto a legacy product — the products are agent systems.
Our products span multi-agent marketing content generation, real-time AI voice intake, and pre- and post-sales intelligence for SMBs — Zylo, IRIS, Ignite, and Ember. They share a common internal pipeline for orchestration, retrieval, tool use, evaluation, and deployment.
This role owns the architecture underneath that pipeline. Not one feature — the shared layer every product depends on: how retrieval is built and measured, how agents reason and hold state and fail safely, which models run where and what happens when one degrades, how output quality is evaluated before it ships, and where the line sits between what a model decides and what deterministic code decides. You’ll set that architecture, build the reusable patterns products inherit, guide the engineers implementing them, and own whether it holds up in production at volume.
It’s a senior individual-contributor role. Two boundaries, stated plainly:
If your GenAI experience is mostly notebooks and proofs of concept that never carried real traffic, this is not the role.
What you’ll own
AI system architecture. Design the systems behind content generation, business intelligence, recommendations, and agentic workflows — including where the boundary sits between what a model decides and what deterministic code decides. Create reusable patterns, prompts, evaluation flows, and orchestration layers that hold up across products rather than being rebuilt per feature.
Agentic workflows. Lead the design of agents that reason, call tools and APIs, manage state, checkpoint, and hand tasks off cleanly. Build the guardrails that keep them grounded and predictable — the hard part is not making an agent act, it’s making it fail safely and legibly when it’s wrong.
Retrieval and model strategy. Design RAG pipelines end to end: chunking, embeddings, metadata, reranking, and retrieval evaluation that actually measures whether the right context arrived. Own model selection across ecosystems on the axes that matter — cost, latency, accuracy, reliability — and define the fallback and model-switching behavior for when a provider degrades.
Evaluation and quality. Stand up the evaluation frameworks that decide whether an AI output ships: output quality, edit ratio, hallucination rate, schema adherence, latency, failure rate, inference cost. Build regression testing so a prompt or model change can’t silently break production. Turn “the output feels off” into a measured threshold.
Production readiness. Partner with engineering and platform teams so systems are deployable, observable, and maintainable. Package services (Python, FastAPI/Flask, Docker) when it’s the fastest path, and diagnose the production failure modes specific to AI — rate limits, cost spikes, model failures, degraded output — rather than escalating them blind.
Technical leadership. Mentor GenAI engineers, review designs and prompts and workflows and evals, and translate product requirements into architecture with measurable acceptance criteria. Communicate the tradeoffs to engineers and to leadership in language each can act on.
What we’re looking for
Required
Preferred
What we care most about
What this role is not
What we offer
Apply if you’ve shipped GenAI systems that carried real traffic and you want to own the architecture of what comes next.