Not currently open. Kept for reference โ do not submit against this role.
๐งโ๐งโ๐งโ๐ง Team & Environment Context
- This is not a research role and not a pure people-management role. We want a builder who has done the hard technical work themselves โ and can also lead a small team of AI engineers from the front.
- The mission: own the "brain" and architecture of our core product at EverAI / Candy.ai โ the conversational AI experience that is central to everything we do.
- The Lead will not have time to learn hard skills on the job. They need to arrive with hands-on fine-tuning experience already behind them.
- The Content: Primarily Candy.ai. Candidates must be 100% comfortable working with NSFW/explicit training data and building moderation for it. NSFW background is a plus, not a requirement.
- Remote, fast-paced scale-up environment. No bureaucracy, no long feedback loops.
Critical Context: The most important thing is not the scale of tokens or infrastructure โ it's whether the AI/chatbot product they worked on was core to the business. A candidate who built a central, user-facing conversational product at a company of 500K+ users is more relevant than someone who optimized inference at a large company where AI was a side feature.
๐ฏ Calibration Anchors & Target Profiles
We are looking for candidates where the conversational AI or chatbot was the product โ not a feature.
- Consumer AI & Companion Platforms
- Creative & Roleplay AI
- High-Scale Conversational Products
- LLM-first startups where the chat experience is the core value proposition
Minimum size guidance: the AI should have been the central revenue-generating and/or impactful feature of the business (it does not matter if it's B2B or B2C).
โ๏ธ Non-Negotiables
- Hands-on fine-tuning experience โ must have done it themselves, for external user-facing products (not internal tooling). Key question: what did you fine-tune, and why? Fine-tuning must represent ~60%+ of their AI experience (agentic AI / RAG / prompt engineering alone is not a fit). Post-training experience is mandatory; pre-training is a nice-to-have.
- The AI product was core / critical to the business โ not a side project, not an internal tool. It's about ownership, impact and risk. Scope includes LLMs and Computer Vision โ both are valid.
- Leadership experience is a must โ being senior is not enough. Minimum 1โ2 years of hands-on leadership experience required. Must still be technically hands-on.
- Strong English communication skills โ this role requires clear technical and cross-functional communication.
- NSFW comfort โ must be comfortable with explicit training data and content moderation in this context.
Good Keywords to Look out for
Models & frameworks โ core: Mistral ยท Llama ยท DeepSeek ยท Qwen ยท Hugging Face
Production & inference โ strong signal: vLLM ยท PyTorch ยท inference optimization ยท open-source LLM
Fine-tuning techniques โ high signal: SFT ยท RLHF ยท DPO ยท LoRA ยท QLoRA ยท PEFT
Fine-tuning frameworks โ niche, very high signal: Axolotl ยท Unsloth ยท LitGPT
๐ Evaluation Framework
1. Fine-Tuning & Model Alignment โ 40%
Has personally fine-tuned models for external, user-facing products. Understands SFT, RLHF, and DPO not just in theory but from direct experience. Can speak concretely to what they fine-tuned, why, and what the outcome was. Familiar with data engineering โ curating, cleaning, and synthetic data generation.
Hard Filters: Fine-tuning experience limited to internal tooling or demos; cannot describe a specific fine-tuning project and its business impact; prompt engineering only.
2. Technical Leadership โ 25%
Player/coach profile. Sets technical direction and owns the architecture while remaining hands-on in the codebase. Has led small teams of senior engineers, conducted code reviews, and maintained a technical roadmap. Does not delegate everything โ writes code.
Hard Filters: Pure managers who no longer code; pure ICs with no leadership experience.
3. Production Engineering โ 25%
Has shipped conversational AI to real users at scale. Comfortable with the modern LLM stack: PyTorch, Hugging Face, vLLM or equivalent inference engines. Understands latency, throughput, and cost trade-offs in production โ not just in benchmarks.
Hard Filters: Academic or research-only profiles; no production deployment experience.
4. Content Moderation Architecture โ 10%
Understands how to build nuanced, context-aware moderation โ not binary safe/unsafe filters. Knows the difference between lazy safety filtering and intelligent classifier-based moderation. NSFW experience is a plus, not a hard requirement.
Hard Filters: Strong ideological objection to working with NSFW content or explicit training data.
๐ฌ Pre-Screen Questions
Use these questions during your screening call. Include a summary of the candidate's answers in your submission notes in Ashby.
- What is your current situation, and why are you considering leaving?
- Why did you leave your last 2 workplaces?
- Are you comfortable working on a product that involves NSFW content โ including working with explicit training data and building moderation for it?
- Tell me about a model you fine-tuned for a live, user-facing product. What was the goal, what did you actually do, and what was the outcome?
- Was the AI/chatbot product you worked on the core of the business, or a feature within a larger product?
- What is your philosophy on model alignment (SFT vs. RLHF/DPO) for a creative, roleplay-centric product? How do you balance safety with steerability?
- Which open-source models (e.g. Llama 3, Mistral) have you been working with recently? What are their strengths and limitations for conversational use cases?
โ๏ธ FAQs
Technical
1. Model Strategy & Ownership
Are you building proprietary models, or fine-tuning open-source / API models? We fine-tune open-source models and also use some APIs. The majority of our traffic โ roughly 80โ90% โ runs on self-hosted fine-tuned open-source models.
What % of usage is OpenAI / Anthropic APIs vs self-hosted? 10โ20% third-party API, 80โ90% self-hosted fine-tuned models.
What is your long-term model strategy? Our strategy is to stay on top of the frontier of open-source models and fine-tune them to match our user preferences. We've tested very large open-source models, and our smaller in-house fine-tuned model still outperforms them on our roleplay evaluation metrics. Roleplay is a highly specific use case, and we believe owning the fine-tuning loop is a durable edge.
How differentiated is your model vs competitors? On our internal roleplay evals, our fine-tuned smaller model outperforms significantly larger out-of-the-box open-source and closed-source models. We haven't run a formal head-to-head benchmark against competitor products, but our qualitative signal โ user retention, session length, and like rates โ suggests a meaningful quality gap in our favor for this specific use case.
2. Infrastructure & Scale
What is the infra stack (cloud, orchestration, serving)? Everything runs on GKE, managed by a 5-person DevOps team.
How are you handling your token volume technically? We serve multi-billion tokens/day of combined input/output on GKE, running across a mix of GPU types.
Latency benchmarks? Kept private. We target sub-2 second P50 latency on self-hosted models for the chat experience.
Cost per request? Kept private, but cost-efficiency is a core optimization axis given our scale.
GPU setup / budget? Budget kept private. Stack runs on modern datacenter GPUs on GKE, e.g. H100s/A100s, RTX PRO.
Any bottlenecks today (infra vs model vs product)? On the infra side, the main challenge is serving larger models (32B+) at efficient speed and cost. But the bigger challenge is on the modeling side โ training a state-of-the-art roleplay model. That's precisely where this hire comes in.
3. Data & Training Pipelines
What is the source of training data? A mix of sources including curated internal data and synthetic generation. Specific breakdowns are kept private, but the candidate would have meaningful data to work with from day one.
User conversations? Synthetic generation? Yes, synthetic generation is part of the mix. Details on both kept private.
How do you handle data labeling? We have an internal process in place. Specifics kept private, but scaling and improving the labeling pipeline is one of the areas we expect this hire to help evaluate.
Preference data for RLHF/DPO? Yes, we run preference-based post-training. The maturity of this pipeline and how to scale it is one of the concrete areas the new hire would help shape.
What tooling exists for experiment tracking / evaluation? We have tooling in place for both. Specifics kept private, but improving experimentation velocity and evaluation maturity are part of the scope.
How mature is your eval framework (offline + online)? Offline: solid and roleplay-specific. Online: earlier stage โ this is explicitly part of the 90-day scope.
4. Moderation & Safety
How does EverGuard actually work? It's a hybrid approach combining model-based classifiers and rules. Specific architecture details are kept private.
What % of engineering effort is moderation vs core product? Around 25% moderation, 75% core product.
What are the hardest unsolved moderation problems today? The hardest ongoing challenges are around context-dependent nuance โ staying permissive enough for legitimate creative/roleplay use cases while catching genuinely harmful patterns, and keeping false-positive rates low so the experience doesn't feel over-censored. These are active areas of work rather than blockers.
Any regulatory exposure? (EU AI Act risk, etc.) No material exposure at this time. We monitor the regulatory landscape proactively.
5. Architecture & Engineering Reality
What does the current system architecture look like end-to-end? Specifics kept private, but the stack is GKE-based with self-hosted inference, a fine-tuning pipeline, moderation layer, and the product surface.
Where is the biggest technical debt? Our training pipeline could be more automated and easier to iterate on โ today, running a new fine-tuning cycle takes more manual orchestration than we'd like. Our online eval infrastructure is also less mature than our offline side. Both are explicitly part of the scope for this hire.
How mature is CI/CD? Monitoring? Incident handling? CI/CD and monitoring are functional but have room to improve โ solid foundation on core infra, less mature on model-level observability. Incident handling is strong, with clear on-call rotations and a good uptime track record.
How often are models deployed to production? Every 2โ3 months depending on cycles and team focus. A faster release cadence is something we want to move toward.
6. Product Performance & Metrics
What are the core success metrics for AI quality? Retention, revenue per user, session length, and like rates.
What are current failure modes in user experience? The biggest one is personalization โ we serve a user base with very different preferences (short messages vs long narrative, different content styles, different roleplay patterns) and our current model doesn't yet adapt dynamically enough to each user. There are also the usual roleplay-model failure modes: occasional consistency drift over long sessions, and edge cases where the model doesn't quite hit the desired tone.
Where is the biggest gap โ Intelligence / Latency / Safety? Primarily intelligence and personalization. Latency and safety are in a good place today, though we keep investing in both.
7. Immediate Scope for This Hire
What are the top 3 problems this person must solve in 90 days? (1) Deeply understand our current training pipeline, identify bottlenecks, and propose concrete improvements. (2) Work closely with the engineering team to upgrade our fine-tuning pipeline, and partner with the CTO and leadership to define the success metrics and KPIs that define what "a better model" actually means for us. (3) Bring product-minded ideas for how to push the roleplay chat experience to the next level, in close collaboration with the product team. By the end of 90 days, the expectation is a clear plan for what a better model looks like, plus concrete initiatives to improve the in-chat user experience.
What is currently blocking growth that this hire unlocks? Two things. First, it frees up significant CTO bandwidth. Second, and more importantly, it brings dedicated ML leadership into the room. Our CTO's background is in web engineering, and we want a peer who owns the ML depth โ someone who can partner with him on technical direction and push our AI engineers on the modeling decisions that will define the next phase of the product.
How much freedom vs alignment with CTO? Trust is earned, and ownership is granted as soon as trust is established. We value people who take real ownership and move fast.
Non-Technical / Commercial
1. Business Model & Monetisation
How do you make money today? Subscriptions and token-based monetization.
What % of users are active (DAU/MAU) / paying? ARPU / LTV? All specific numbers kept private at this stage. Happy to share more context with candidates who reach the later interview stages under NDA.
2. Growth Quality
"50M users" โ how many are real vs bots? We have captcha plus email confirmation, so bot contamination is low. The vast majority are real users.
Retained after 30 days? Kept private. Happy to discuss ranges with candidates in later stages.
What channels drove growth? (viral vs paid) Kept private at this stage.
3. Competitive Landscape
Who do you see as direct competitors? Character.ai and Replika in the mainstream roleplay space. In the NSFW-oriented segment, Ourdream.ai and Lovescape. We view ourselves as differentiated on model quality for the roleplay use case specifically.
4. Team & Hiring Plan
What is the target team size in 6โ12 months? Depends on needs and the business trajectory โ we hire based on clear impact, not headcount targets.
Will this hire become Head of AI, or stay as senior IC / player-coach? We're still a small company, so everything is possible. Anyone who shows extreme dedication, motivation, and business impact can grow into larger responsibilities. The initial framing is senior IC / player-coach with a clear path upward.
What other hires are planned around them? None currently planned for the next 3โ6 months, unless a specific project with high business impact justifies it.
5. Culture Reality Check
What does "high performance" actually mean here? Ownership, curiosity, strong motivation, good attitude, and measurable business impact.
Hours / output expectations? We optimize for output and impact, not hours. That said, we move fast and expect people to match that energy.
Any recent attrition? No material attrition issues.
How are disagreements handled? A quick 1:1 or a focused call with the relevant stakeholders to align on direction. We value speed of decision-making.
6. Risk & Controversy
Past platform bans? Payment provider issues? PR issues due to NSFW positioning? None on all three.
How do you mitigate platform risk (Apple/Google policies)? No issues today. Our positioning and compliance posture keep us clear of platform problems.
7. Compensation Clarity
Equity = $80k total or annual? Total. Package is competitive and structured for the long term.
Flexibility on base? The package is competitive and we can structure it for the right profile. Specifics to be discussed with your TA contact.
Any upside tied to growth milestones? Not part of the current structure.