ai evaluation engineer
10,000 ofert pracy ai evaluation engineer w Polska. Codziennie aktualizowane oferty.
-
Senior AI Evaluation Engineer for Frontier Models
19 godzin temu
Warszawa, Województwo mazowieckie, Polska NVIDIA Pełny etat 390 000 zł - 650 000 zł UmowaNVIDIA is seeking senior engineers to pioneer evaluation methodologies for state-of-the-art AI models, including LLMs, vision, and agentic systems. You will build and scale experiments on enterprise-grade GPU clusters to influence model releases and roadmap decisions.Collaborating with model research, training, and product teams, you will design evaluation...
-
Freelance Agent Evaluation Engineer Premium
18 godzin temu
Poland Mindrift Zdalnie Pełny etat 40 USD Na stałeDescriptionPlease submit your CV in English and indicate your level of English proficiency.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.We're building a dataset to evaluate AI coding agents -...
-
Software Engineer III, Evaluation Platform
19 godzin temu
Warszawa, Województwo mazowieckie, Polska Box Pełny etat 180 000 zł - 230 000 zł UmowaSoftware Engineer III, Evaluation Platform Warsaw, Poland What is Box? Box (NYSE:BOX) is the leader in Intelligent Content Management. Our platform enables organizations to fuel collaboration, manage the entire content lifecycle, secure critical content, and transform business workflows with enterprise AI. We help companies thrive in the new AI-first era of...
-
Software Engineer III, Evaluation Platform Premium
3 dni temu
Warsaw, Polska Box Pełny etatWhat is Box? Box (NYSE:BOX) is the leader in Intelligent Content Management. Our platform enables organizations to fuel collaboration, manage the entire content lifecycle, secure critical content, and transform business workflows with enterprise AI. We help companies thrive in the new AI-first era of business. Founded in 2005, Box simplifies work for...
-
Senior Data Scientist – AI/LLM Evaluation
2 godzin temu
mazowieckie, mazowieckie, Polska Pełny etatWe are looking for a Senior Data Scientist to join an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation.The project focuses on developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows. The role involves designing...
-
Technical Lead, AI Evaluation
2 dni temu
Kraków, Województwo małopolskie, Polska InPost UK Pełny etat 320 000 zł - 420 000 zł UmowaInPost Group, the team behind Von Halsky, seeks a Technical Lead, Senior AI Engineer to own the Evaluations Platform and guide the next phase of production-grade AI tooling. You will oversee the LLM-judge, curate eval datasets, and drive measurable improvements in conversation quality.You will mentor engineers, define the technical direction, and collaborate...
-
Warszawa, Województwo mazowieckie, Polska Mercor Zdalnie Pełny etatMercor partners with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing structured assessments on infrastructure tasks and model outputs.The role focuses on real-world engineering workflows, cloud platforms, Kubernetes, and automation. The position requires 2+ years in...
-
AI Safety Specialist
4 dni temu
Warszawa, Województwo mazowieckie, Polska Obsidian Pełny etat 180 000 zł - 300 000 zł UmowaWe are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback.ResponsibilitiesEvaluate...
-
Freelance Agent Evaluation Engineer
4 dni temu
, Polska Mindrift Pełny etatPlease submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. We're building a dataset to evaluate AI coding agents - how well a...
-
Senior Data Scientist — AI Evaluation Premium
18 godzin temu
Krakow, Polska Finom Zdalnie Pełny etatAbout Finom Finom is a European tech startup headquartered in Amsterdam, and we’re on a journey towards revolutionizing the financial landscape for entrepreneurs worldwide. Our mission is to develop an all-in-one financial B2B solution that integrates banking functions, accounting, financial management, and invoicing into a seamless, mobile-first...
-
Senior Data Annotator
19 godzin temu
Warszawa, Województwo mazowieckie, Polska Rex.zone Pełny etat 60 000 zł - 90 000 zł UmowaRex.zone in Warsaw (Remote) is seeking senior data annotators to create and validate high-quality training data for AI/ML systems. You will support LLM training pipelines through data labeling, QA evaluation, prompt evaluation, and RLHF-style ranking to improve model behavior and evaluation quality.You will annotate NLP, vision, and safety datasets, perform...
-
Senior AI Evaluation Architect for Frontier Models
19 godzin temu
, Polska NVIDIA Corporation Zdalnie Pełny etat 375 000 zł - 650 000 zł UmowaNVIDIA Corporation is seeking senior engineers to design evaluation environments for cutting-edge AI models, including LLMs, multimodal systems, and agents. You will build infrastructure, pursue novel evaluation methods, and collaborate with research and product teams to influence model releases.The role emphasizes strong statistical analysis, scalable...
-
Senior AI Data Scientist — LLM Metrics
2 dni temu
Bydgoszcz, Województwo kujawsko-pomorskie, Polska Knowit Pełny etat 180 000 zł - 240 000 zł UmowaKnowit w Polsce poszukuje Salesforce Marketing Cloud Engineera na poziomie Mid/Senior, który wesprze rozwój skalowalnych rozwiązań marketingowych na globalnej platformie e-commerce.Twoje zadania obejmą projektowanie metryk ewaluacyjnych, tworzenie danych ground truth oraz ocenę trafności wyników i jakości ewaluacji, a także rozwijanie metod ocen z...
-
Senior Data Scientist — AI/LLM Evaluation
2 dni temu
, Polska Square One Resources Zdalnie Pełny etat 180 000 zł - 240 000 zł UmowaSquare One Poland is seeking a Senior Data Scientist to join an international project focused on AI, Large Language Models (LLMs), and AI evaluation. You will design evaluation metrics, build ground-truth datasets, analyze model performance, and advance automated and human-feedback-based evaluation approaches.The role emphasizes evaluating AI agents and...
-
Data Scientist: AI Evaluation
2 dni temu
Województwo kujawsko-pomorskie, Województwo kujawsko-pomorskie, Polska Knowit Pełny etat 180 000 zł - 260 000 zł UmowaKnowit Connectivity w Polsce poszukuje specjalisty ds. ewaluacji AI do pracy nad oceną jakości systemów AI i metryk ewaluacyjnych. Do zadań należy projektowanie i walidacja metryk, tworzenie zestawów ground truth oraz analiza trafności i błędów.Doświadczenie w ewaluacji AI i modelach LLM będzie kluczowe. Oferujemy pracę w dynamicznym zespole,...
-
Freelance Agent Evaluation Engineer Premium
4 dni temu
Jordanów, Lesser Poland Voivodeship Mindrift Zdalnie Pół etatu 50 USD Na stałePlease submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. We're building a dataset to evaluate AI coding agents - how well a...
-
Freelance Agent Evaluation Engineer Premium
4 dni temu
Gdańsk, Pomeranian Voivodeship, Polska Mindrift Zdalnie Pół etatu 40 USD Na stałePlease submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. We're building a dataset to evaluate AI coding agents - how well a...
-
Technical Lead — AI Evaluation
4 dni temu
Kraków, Województwo małopolskie, Polska InPost Pełny etat 380 000 zł - 680 000 zł UmowaInPost seeks a Technical Lead, Senior AI Engineer to own the Evaluations Platform for VonHalsky, InPost's conversational AI assistant. You will define technical direction, mentor engineers, and ensure production-grade evals before features reach users.You will drive the LLM-judge framework, curate real and golden data, and maintain robust diagnostics while...
-
Senior Deep Learning Engineer, Accuracy Evaluation Premium
18 godzin temu
Switzerland, Remote, Polska NVIDIA Zdalnie Pełny etat 375 000 zł - 650 000 zł UmowaWe are seeking senior engineers to pioneer new methodologies for accurately assessing the performance and capabilities of ground-breaking deep learning models, including LLMs, RAG, agents, and vision models. You will collaborate across the organization to bring the latest flagship models from our community and partners to life. This role offers an...
-
Senior Deep Learning Engineer, Accuracy Evaluation
19 godzin temu
Warszawa, Województwo mazowieckie, Polska NVIDIA Pełny etat 390 000 zł - 650 000 zł UmowaWe are seeking senior engineers to pioneer new methodologies for accurately assessing the performance and capabilities of ground-breaking deep learning models, including LLMs, RAG, agents, and vision models. You will collaborate across the organization to bring the latest flagship models from our community and partners to life. This role offers an...