carnaby fox is actively hiring for 2 AI and machine learning positions, concentrated in Prompt Engineer (1) and AI Agent Developer (1) roles. Positions are based in San Francisco, CA, US, Margaretville, NY, US. The most frequently requested skills across these postings are Prompt Engineering, Python, Rag. Mid-level roles account for 50% of openings.
Skills & Technologies
Locations
San Francisco, CA, US, Margaretville, NY, US
Hiring by Role Category
Open Positions (2)
What carnaby fox's hiring tells you
With 2 active AI role(s), this company is in the early exploration phase. That can mean either a pilot project being staffed up or a small embedded AI function inside a larger team. Worth investigating directly: ask the recruiter how the AI work is funded and who it reports to. Compensation is not disclosed in postings, which is increasingly out of step with how AI talent expects to be hired.
The skill mix here leans toward Prompt Engineering in Prompt Engineer roles. That is a clue about what carnaby fox is building: teams hire for the work in front of them, not the work they wish they were doing.
Questions worth asking in the carnaby fox interview loop
The signals above come from public job postings. The signals you actually need come from the conversation. A few questions calibrated to this company's tier:
- Is this AI work funded for at least 18 months, or is it tied to a specific project deadline?
- Will I be the only person doing this, or are there others I will collaborate with day to day?
- What does success look like at six months? At eighteen months?
carnaby fox AI and ML Hiring
carnaby fox has 2 active AI and ML roles in our dataset. Open positions span Prompt Engineer, AI Agent Developer. Roles are based in San Francisco, CA, US, Margaretville, NY, US.
Salary Benchmarks
The market median for AI roles is $217,500. Prompt Engineer roles pay a median of $140,000 across the market. AI Agent Developer roles pay a median of $238,500 across the market. Top-quartile AI compensation starts at $272,100.
Skills carnaby fox Looks For
The core requirement is deep LLM experience: prompt design, RAG architectures, and evaluation methodology. Python is table stakes. Many roles also want experience with specific providers like OpenAI, Anthropic, or open-source models. Understanding tokenization, context windows, and the practical differences between model families (reasoning ability, instruction following, output format compliance) separates strong candidates from the crowd.
Evaluation skills are becoming the differentiator. Can you design a rubric that measures output quality? Can you build automated evaluation pipelines? Do you understand when to use human evaluation vs. LLM-as-judge vs. deterministic checks? Companies are moving past 'vibes-based' prompt testing and want engineers who bring measurement discipline.
AI Role Categories
Prompt Engineer
Prompt Engineers design, test, and optimize interactions with large language models. They build evaluation frameworks, craft system prompts, and develop techniques like chain-of-thought and few-shot learning to get consistent, reliable outputs. The role emerged alongside the GPT-3 era and has matured into a legitimate engineering discipline, not the 'just talk to the AI' job that early skeptics dismissed.
The core requirement is deep LLM experience: prompt design, RAG architectures, and evaluation methodology. Python is table stakes. Many roles also want experience with specific providers like OpenAI, Anthropic, or open-source models. Understanding tokenization, context windows, and the practical differences between model families (reasoning ability, instruction following, output format compliance) separates strong candidates from the crowd.
Market compensation for Prompt Engineer roles: $140,000 median across 11 positions with disclosed pay.
AI Agent Developer
AI Agent Developers build autonomous systems that can reason, plan, and take actions. They design multi-step workflows, tool-use frameworks, and orchestration layers that let LLMs interact with external systems. This is the frontier of applied AI engineering.
Deep experience with LLM APIs and agent frameworks (LangChain, CrewAI, AutoGen). Strong understanding of prompt engineering, function calling, and error handling for non-deterministic systems. Python is standard. Experience with orchestration patterns, state management, and workflow engines adds significant value.
Market compensation for AI Agent Developer roles: $238,500 median across 58 positions with disclosed pay.
The AI Job Market Today
The AI job market spans 3,708 open positions across 16 role categories. The largest categories by volume: AI/ML Engineer (2,605), Data Scientist (310), AI Software Engineer (259). These three account for the majority of open positions, though smaller categories often have higher per-role compensation because of specialized skill requirements.
The seniority mix tells a story about where AI teams are in their maturity. Entry-level roles (102) are outnumbered by mid-level (1,705) and senior (1,469) positions, reflecting that most companies are past the 'build a team from scratch' phase and need experienced engineers who can ship production systems. Leadership roles (Director, VP, C-Level) total 432 positions, representing the bottleneck between technical execution and organizational strategy.
Remote work availability sits at 14% of all AI roles (508 positions), with 3,180 requiring on-site or hybrid attendance. The remote share has stabilized after the post-pandemic correction. Senior and specialized roles (Research Scientist, ML Architect) are more likely to be remote-eligible than entry-level positions, partly because experienced hires have more negotiating power and partly because these roles require less hands-on mentorship.
AI compensation is structured in clear tiers. The market median sits at $217,500. Top-quartile roles start at $272,100, and the 90th percentile reaches $325,000. These figures include base salary with disclosed compensation. Total compensation (including equity, bonuses, and sign-on) runs 20-40% higher at companies that offer those components.
Category matters for compensation. AI Safety roles lead at $300,000 median, while Prompt Engineer roles sit at $140,000. The spread between highest and lowest-paying categories reflects the premium on specialized technical skills versus broader analytical roles.
The most in-demand skills across all AI postings: Python (1,890 postings), Aws (1,103 postings), Azure (877 postings), Rag (855 postings), Gcp (631 postings), Prompt Engineering (560 postings), Pytorch (545 postings), Claude (498 postings). Python dominates, appearing in the vast majority of role descriptions regardless of category. Cloud platform experience (AWS, GCP, Azure) is the second most common requirement. The newer entrants to the top skills list (RAG, vector databases, LLM APIs) reflect the shift from traditional ML toward generative AI applications.
AI Hiring Overview
The AI job market has 3,708 open positions tracked in our dataset. By seniority: 102 entry-level, 1,705 mid-level, 1,469 senior, and 432 leadership roles (Director, VP, C-Level). Remote roles make up 14% of the market (508 positions). The remaining 3,180 roles require on-site or hybrid attendance.
The market median for AI roles is $217,500. Top-quartile compensation starts at $272,100. The 90th percentile reaches $325,000. Highest-paying categories: AI Safety ($300,000 median, 21 roles); Research Engineer ($280,000 median, 147 roles); AI Architect ($254,798 median, 67 roles).
Prompt engineering roles are still growing but the market is maturing. Early roles were broad and experimental. Now, companies know what they want: someone who can systematically improve LLM output quality, reduce costs by optimizing token usage, and build evaluation infrastructure. The roles that survive will be the ones that look more like engineering than copywriting.
What to Expect in Interviews
Interviews focus on evaluation methodology and systematic thinking. You'll likely be asked to design a prompt for a specific use case, explain how you'd measure output quality, and walk through how you'd debug a prompt that works 90% of the time but fails on edge cases. Expect to discuss tokenization, context window management, and the tradeoffs between different prompting strategies (few-shot vs. chain-of-thought vs. tool use).
When evaluating opportunities: Strong postings specify the LLM use cases (summarization, extraction, classification, generation), the evaluation methodology they expect, and the production environment. Weak postings just say 'prompt engineering experience' without context. Look for companies that mention evaluation frameworks and production deployment.
Frequently Asked Questions
Frequently Asked Questions
Related Resources
Get Weekly AI Career Intelligence
Salary data, skills demand, and market signals from 16,000+ AI job postings. Every Monday.