AI Research Engineer (Pre-training - LLM & Multi-Modal) - 100% Remote Worldwide

Tether
13 days
Remote, Worldwide
Engineer Research AI Blockchain LLM PyTorch Fintech Machine Learning Remote Scalability Architecture Design Worldwide Multi-Modal Pre-training Hugging Face NVIDIA GPUs Distributed Training Transformer Architectures NLP AI R&D Tokenizers Cross-Modal Alignment Data Curation Model Optimization Debugging Computer Science
Join Tether and shape the future of digital finance. Tether is pioneering a global financial revolution, empowering businesses with cutting-edge solutions to integrate reserve-backed tokens across blockchains. Our innovative product suite includes the world’s most trusted stablecoin, USDT, and pioneering digital asset tokenization services (Tether Finance). We also drive sustainable growth through energy solutions for Bitcoin mining (Tether Power), fuel breakthroughs in AI and peer-to-peer technology with solutions like KEET (Tether Data), democratize digital learning (Tether Education), and push boundaries at the intersection of technology and human potential (Tether Evolution). Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about fintech and making a significant impact, this is an opportunity to collaborate with bright minds, pushing boundaries and setting new industry standards. We value excellent English communication skills and invite you to be part of the most innovative platform on the planet. As a member of the AI model team, you will drive innovation in architecture development for cutting-edge models of various scales, including small, large, and multi-modal systems, aiming to enhance intelligence, improve efficiency, and introduce new capabilities. You will apply deep expertise in Large Language Model (LLM) and Multi-Modal architectures, a strong grasp of pre-training optimization, and a hands-on, research-driven approach. Your mission involves exploring and implementing novel techniques and algorithms that lead to groundbreaking advancements: including multi-modal data curation and alignment, strengthening baselines, and identifying and resolving existing pre-training bottlenecks to push the limits of cross-modal AI performance. Responsibilities include: conducting foundational large-scale pre-training for LLMs and Multi-Modal models (integrating text, vision, audio, or other modalities) on large, distributed servers with multi-nodes and thousands of NVIDIA GPUs; designing, prototyping, and scaling innovative architectures, tokenizers, and cross-modal alignment layers to enhance model intelligence and multi-modal understanding; sourcing, filtering, and curating massive-scale textual and multi-modal datasets, establishing robust data pipelines for efficient pre-training; independently and collaboratively executing experiments, analyzing results, and refining training methodologies for optimal performance and token efficiency; investigating, debugging, and eliminating bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during long training runs; and contributing to the advancement of distributed training systems to ensure seamless scalability and hardware efficiency on target platforms.