Embodied AI
Vision–Language–Action systems that ground language in motion — modular stacks, training loops, and sim-to-real plumbing.
We treat perception, language, planning, and control as specialized modules coordinated through a World Model rather than one monolithic transformer. Kinetic and DifoTrain are the product surface of this line.