Senior Applied AI Evaluation Engineer - RAN Ticket Intelligence
Parallel Wireless
Apply to this jobWe are looking for a hands-on Senior Applied AI Evaluation Engineer to improve how our RAN R&D organization analyzes engineering tickets, supports Root-Cause Analysis, recommends ownership, and learns from resolved cases.
You will build rigorous evaluation datasets, establish meaningful baselines, compare internal and approved external tools, analyze failure modes, and prototype improvements across retrieval, classification, prompting, agent workflows, and model selection.
This role is focused on measurable, evidence-based improvement rather than AI demonstrations. You will assess whether AI-generated conclusions are accurate, grounded in evidence, appropriately calibrated, and useful to engineering teams.
As a secondary area of focus, you will analyze engineering workflows at case and team level to identify bottlenecks, handoffs, dependencies, and opportunities for process improvement.
What you will do:
- Define high-value RAN ticket-intelligence use cases, acceptance criteria, evaluation metrics, and quality guardrails.
- Build and maintain representative, versioned evaluation datasets using resolved tickets, Root-Cause Analysis, logs, test evidence, code changes, reviews, reassignment history, and outcomes.
- Establish current-tool and non-AI baselines before evaluating new LLM, RAG, search, or agent-based approaches.
- Evaluate approved internal, commercial, local, and open-source solutions using secure and reproducible data-handling processes.
- Measure retrieval quality, groundedness, diagnosis accuracy, citation support, routing recommendations, calibration, abstention, latency, cost, and human effort.
- Design held-out, time-based, edge, and adversarial test cases while preventing data leakage and future-outcome contamination.
- Analyze failures and turn incorrect conclusions, misrouting, unsupported claims, and missed evidence into prioritized improvements.
- Prototype improvements in search, metadata, context construction, prompting, reranking, classification, agent workflows, and model selection.
- Develop reusable evaluation pipelines, tools, services, APIs, dashboards, or documented workflows.
- Work closely with AI, RAN, QA, System Integration, Release, Field, data, and engineering teams to review results and support evidence-based decisions.
What you should have:
- BSc or MSc in Computer Science, Data Science, Machine Learning, Statistics, Electrical Engineering, or a related field, or equivalent practical experience.
- 5+ years of hands-on experience in applied machine learning, data science, search, natural-language processing, analytics engineering, or AI-enabled software systems.
- Recent experience evaluating LLM, RAG, search, or agent systems using representative datasets, task-specific metrics, human review, failure analysis, and regression testing.
- Strong Python and SQL skills, with experience building maintainable data pipelines, experiment workflows, services, or analytical tools.
- Practical experience with several areas such as information retrieval, embeddings, hybrid search, reranking, classification, structured outputs, tool calling, or common LLM failure modes.
- Strong statistical judgment, including sampling, leakage prevention, uncertainty, calibration, precision and recall, temporal drift, and controlled comparison of competing approaches.
- Ability to work with semi-structured engineering data from issue-tracking systems, source control, code reviews, continuous integration, logs, dashboards, and test systems.
- Clear communication skills and the ability to explain results, limitations, and tradeoffs to technical and business stakeholders.
Preferred qualifications:
- Knowledge of LTE, 5G NR, Open RAN, telecom-support workflows, or demonstrated ability to learn a technically complex domain through close collaboration with subject-matter experts.
- Experience with enterprise search, RAG evaluation, knowledge graphs, process mining, anomaly detection, or graph-based analysis.
- Familiarity with Jira, Git or Bitbucket, CI/CD telemetry, software-delivery analytics, evaluation frameworks, experiment tracking, or data versioning.
- Experience working with open-weight LLMs, commercial model APIs, local inference, proprietary code, customer logs, or access-controlled engineering data.
- Experience handling privacy-sensitive or regulated data in secure enterprise environments.
Summary
Evaluate AI systems, build datasets, analyze failures, prototype improvements for RAN tickets.
Job title
Senior Applied AI Evaluation Engineer
Experience level
5+ years
Minimum experience
5+ years exp
Industry
telecommunications
Location requirements
Kfar Saba, no remote work allowed
Salary
Not specified
Visa sponsorship
H-1B sponsor history
Management role
No
Required skills
Preferred skills
Specializations
Structured locations inferred from the posting.
Kefar Sava, Israel