Job Description
THE ORGANIZATION
The Alliance of Bioversity International & CIAT delivers research-based solutions that harness agricultural biodiversity and sustainably transform food systems to improve people’s lives. Alliance solutions address the global crisis of malnutrition, climate change, biodiversity loss, and environmental degradation.
The Alliance works with local, national, and multinational partners across Sub Saharan Africa, Latin America and the Caribbean and Asia, and with public and private sectors. The Alliance is part of CGIAR, a global research partnership for a food-secure future, dedicated to reducing poverty, enhancing food and nutrition security, and improving natural resources and ecosystem services.
About the position
The Senior Research Associate - Data Management and Analytics will plan and oversee data management and analytics operations across the NDIZI and Google.org projects, developing and implementing data strategies, monitoring progress against established milestones, and organizing data collection, processing, quality assurance, and reporting services across multiple field sites and partner institutions. This role leads planning of the project’s data infrastructure, ensures data quality and governance, and prepares high-quality datasets for AI model development to accelerate crop improvement and enhance global food security.
Key duties & responsibilities
Data Management
● Oversee the development, implementation, and continuous improvement of the project's data management strategy.
● Manage the organization, integration, and governance of datasets from diverse sources, including field trials, image repositories, partner institutions, and digital platforms.
● Contribute to the development and maintenance of comprehensive data documentation, metadata standards, and data dictionaries to ensure long-term usability and reproducibility.
Data Processing and Analysis:
● Lead the design, development, and optimization of data processing pipelines for machine learning and analytical workflows.
● Manage data storage, transformation, and integration solutions, ensuring scalability, security, and interoperability.
● Guide efforts to convert data between different formats as needed.
● Process multi-modal datasets (images, speech/text, tabular field trial data) to the quality and format standards required for downstream AI/ML workflows.
● Conduct exploratory data analysis and statistical summaries to surface patterns, outliers, and insights informing model development and research reporting.
● Monitor quality review pipeline performance and maintain version controlled, documented data processing workflows to ensure reproducibility and research integrity
Data Quality Control:
● Plan, Design and implement robust data quality control procedures.
● Manage data cleaning, validation, and verification tasks to ensure data accuracy and consistency.
● Identify and report data anomalies and inconsistencies.
● Advise the AI/ML team on data quality standards, known anomalies, and preprocessing requirements, ensuring that training datasets meet the quality thresholds needed for reliable model performance
Data Annotation and Labelling Oversight:
● Provide technical guidance for the annotation and labelling of image and structured data for supervised machine learning tasks.
● Collaborate with domain experts and model developers to define labeling schemas, annotation workflows, and quality assurance protocols.
● Evaluate and recommend tools, platforms, and processes to improve annotation efficiency and accuracy at scale
Capacity Building and Collaboration
● Advice and provide guidance to junior data staff and contributors involved in data collection, management, and processing activities.
● Develop training resources and lead capacity-building initiatives on data management best practices and tools for internal teams and external partners.
● Serve as a key liaison between data, research, and engineering teams to ensure coordinated and efficient data flows that support modelling, analytics, and decision-making.
● Represent the data function in strategic planning meetings, reporting processes, and donor communications.
Requirements
Required qualifications and experience.
• Master's degree in Agricultural Sciences, Computer Science, Data Science, or a related field.
• 3+ years of experience in data management, data analysis, or a related field, demonstrating progressive responsibility in research support contexts.
• Familiarity with data quality control and data processing techniques.
• Experience with data annotation and labelling tools.
• Demonstrated supervisory experience, including overseeing junior data staff, managing workflows, and mentoring team members in data management best practices
• Proficiency with NOSQL and SQL databases for data querying, management, and integration.
• Experience with cloud storage and computing platforms (e.g., AWS S3, Google Cloud Storage) and associated data management workflows.
• Familiarity with ETL/ELT pipeline design and workflow orchestration tools.
• Proficiency with data visualization tools to communicate data insights.
• Experience with version control systems (e.g., Git) and collaborative development workflows for managing data scripts and pipelines.
• Familiarity with FAIR data principles, open data standards, and research data management planning.
• Experience with agricultural and/or plant sciences.
• Proficient programming skills in Python or other relevant languages.
• Familiarity with crop phenotyping data, field trial design, or agricultural data collection workflows.
• Experience with data annotation platforms (e.g., CVAT, LabelStudio, Roboflow) and quality assurance workflows for image and structured data.
• Familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch, Scikit-learn) for data preparation, feature engineering, and pipeline integration.
• Proficiency in data manipulation and analysis.
• Strong attention to detail and accuracy.
• Excellent communication and interpersonal skills.
• Ability to work independently and as part of a team.
Desirable Qualifications:
● Knowledge of speech, text and image processing and analysis techniques.
● Familiarity with metadata standards and best practices.
● Experience in LLM-based analytics workflows
● Experience working in an international agricultural research environment
Benefits
Terms of employment
This is a nationally recruited position based in Arusha, Tanzania. The initial contract will be for one year, subject to a probation period of three months and is renewable depending on performance and availability of resources.
This position is graded at BG07 level, with a minimum basic salary of TZS 3,929,704 in a scale of BG01 to BG14 (BG14 being the highest level according to the Alliance job classification policy). We offer a competitive salary and excellent benefits including but not limited to insurance, retirement plan, staff training and development, paid time off and flexible working arrangements.
The Alliance Bioversity-CIAT is committed to fair, safe, and inclusive workplaces. We believe that diversity powers our innovation, contributes to our excellence, and is critical for our mission. Recruiting and mentoring staff to create an inclusive organization that reflects our global character is a priority. We encourage applicants from all cultures, races, colors, religions, sexes, national or regional origins, ages, disability statuses, sexual orientations, marital status, and gender identities. Female candidates are strongly encouraged to apply.