Deskripsi Pekerjaan
Are you passionate about pushing the boundaries of Artificial Intelligence? Nanyang Technological University (NTU) is seeking a highly skilled AI Engineer (Evaluation) to join our elite research team. In this role, you will be at the forefront of developing rigorous evaluation frameworks and testing pipelines for cutting-edge multilingual large language models (LLMs).
As an AI Engineer, you will bridge the gap between theoretical research and practical application. You will be responsible for designing metrics, conducting benchmark testing, and ensuring our models meet the highest standards of accuracy, safety, and linguistic nuance. You will collaborate with world-class researchers in an innovative, fast-paced environment that fosters academic and technical growth.
If you have a deep interest in NLP, model robustness, and systematic performance measurement, we invite you to help us shape the next generation of multilingual AI technologies.
Tanggung Jawab
- Design and implement comprehensive evaluation pipelines for multilingual AI models.
- Develop automated benchmarks to assess model performance across diverse linguistic tasks.
- Collaborate with research teams to define performance metrics and qualitative assessment criteria.
- Conduct error analysis on model outputs to identify failure modes and edge cases.
- Build and maintain data annotation tools and quality control workflows.
- Document evaluation methodologies and prepare technical reports for publication or internal review.
- Optimize testing infrastructure to support large-scale model validation cycles.
Kualifikasi
- Bachelor’s or Master’s degree in Computer Science, Data Science, AI, or a related quantitative field.
- Proven experience in building AI/ML pipelines, specifically within an NLP or LLM context.
- Strong proficiency in Python and deep learning frameworks such as PyTorch or TensorFlow.
- Solid understanding of evaluation metrics (e.g., BLEU, ROUGE, perplexity, or custom human-alignment metrics).
- Experience working with multilingual datasets and understanding linguistic nuances in AI evaluation.
- Familiarity with cloud computing platforms (AWS, GCP, or Azure) for model training and testing.
- Strong analytical mindset with a passion for meticulous data analysis and problem-solving.