You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
Binance

Senior Evaluation Algorithm Engineer

  • Remote
  • 6+ years

Salary

Not stated

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.

About the Role

In the AI era, large language models are reshaping core business scenarios such as dialogue and trading. Model capability iteration relies on a scientific and trustworthy evaluation system — the "ruler" that measures model quality and guides R&D direction. We are seeking an evaluation expert with an algorithmic background to build LLM evaluation capabilities covering dialogue, financial trading, and other scenarios, using professional evaluation methods to quantify model performance, pinpoint issues, and drive continuous model improvement.

Responsibilities

  • Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.
  • Lead the design and construction of evaluation datasets. Define evaluation dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.
  • Analyze model capability boundaries and failure modes based on evaluation results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.
  • Drive the automation and scaling of evaluation workflows. Build sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.
  • Collaborate with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.

Requirements

  • Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.
  • Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.
  • Familiar with mainstream evaluation methods (human evaluation, model-based automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.
  • Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.
  • Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.
  • Strong business understanding and communication skills, able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration.
  • Bonus Qualifications
  • Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs.
  • Experience building high-quality AI training/evaluation data or data annotation systems.
  • Familiarity with RLHF, reward models, or preference data-related work.
  • Why Binance
  • Shape the future with the world’s leading blockchain ecosystem
  • Collaborate with world-class talent in a user-centric global organization with a flat structure
  • Tackle unique, fast-paced projects with autonomy in an innovative environment
  • Thrive in a results-driven workplace with opportunities for career growth and continuous learning
  • Competitive salary and company benefits
  • Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)
  • Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.
  • By submitting a job application, you confirm that you have read and agree to our Candidate Privacy Notice .

Where you’d work

Fully remote

You can work from

  • Australia

About the company

Binance

  • Industry: FinTech

Also hiring in London, United Kingdom

6 of their 7 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 19 checks run
  • 3 red flags

Still hiring?

13 checks

1 red flag

How crowded?

6 checks

2 red flags

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details12 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • LLMs
Seniority
Senior
Type
Full-time
Industry
FinTech
Specialty
ML
Region
Australia
Show 6 more factsShow less

Role

Category
Data & Analytics
Specialty
ML
Seniority
Senior
Experience
6+ years
Tech stack
  • Python
  • LLMs

Location

Work model
Remote
Region
Australia
Remote from
  • Australia

Compensation

Salary
Salary by agreement
Pay period
Annual

Employment

Type
Full-time

Company

Industry
FinTech

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Data & AI Engineer Intern

    Salary by agreement

    • Hybrid · Paris
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    AI Engineer

    $80,000 - 210,000 / year

    • Office · CA, Austin
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior AI/ML Engineer

    $200,000 - 260,000 / year

    • Office · San Francisco
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Member of Technical Staff, ML Infra

    Salary by agreement

    • Office · San Francisco
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason