AI & ML Papers
34K subscribers
7.37K photos
593 videos
24 files
8.11K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
Media is too big
VIEW IN TELEGRAM
✨Agent S: An Open Agentic Framework that Uses Computers Like a Human

📝 Summary:
Agent S is an open agentic framework enabling autonomous GUI interaction to automate complex tasks. It employs experience-augmented hierarchical planning and an Agent-Computer Interface with MLLMs for enhanced reasoning. Agent S achieves state-of-the-art performance on OSWorld and demonstrates br...

🔹 Publication Date: Published on Oct 10, 2024

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2410.08164
• PDF: https://arxiv.org/pdf/2410.08164
• Github: https://huggingface.co/collections/ranpox/awesome-computer-use-agents

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#AgenticAI #MultimodalAI #HumanComputerInteraction #Automation #AIResearch
✨Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics

📝 Summary:
TIMAR is a new causal framework for 3D conversational head generation. It models dialogue using interleaved audio-visual contexts to predict continuous head dynamics, improving coherence and expressive variability. Experiments show TIMAR significantly reduces errors and improves performance.

🔹 Publication Date: Published on Dec 17

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.15340
• PDF: https://arxiv.org/pdf/2512.15340
• Project Page: https://github.com/CoderChen01/towards-seamleass-interaction/blob/main/README.md
• Github: https://github.com/CoderChen01/towards-seamleass-interaction

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#ConversationalAI #3DAnimation #HumanComputerInteraction #CausalModeling #AI
✨Continual GUI Agents

📝 Summary:
The Continual GUI Agents framework addresses performance degradation in dynamic UI environments. It introduces GUI-Anchoring in Flux GUI-AiF, a reinforcement fine-tuning method with novel anchoring rewards that stabilize learning across shifting UI domains and resolutions, outperforming existing ...

🔹 Publication Date: Published on Jan 28

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.20732
• PDF: https://arxiv.org/pdf/2601.20732

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#ContinualLearning #ReinforcementLearning #AIAgents #HumanComputerInteraction #MachineLearning
This media is not supported in your browser
VIEW IN TELEGRAM
✨Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control

📝 Summary:
This paper introduces a human-centric video world model for extended reality, using tracked head and hand poses for dexterous interaction. This system generates egocentric virtual environments, significantly improving user task performance and perceived control.

🔹 Publication Date: Published on Feb 20

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.18422
• PDF: https://arxiv.org/pdf/2602.18422
• Project Page: https://codeysun.github.io/generated-reality/

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#ExtendedReality #VideoGeneration #HumanComputerInteraction #VirtualEnvironments #AIResearch
❤1
✨How to Take a Memorable Picture? Empowering Users with Actionable Feedback

📝 Summary:
This paper introduces Memorability Feedback MemFeed, a new task providing actionable natural language guidance to improve photo memorability. Their method, MemCoach, uses MLLMs and a teacher-student strategy, demonstrating that memorability can be taught and instructed.

🔹 Publication Date: Published on Feb 25

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.21877
• PDF: https://arxiv.org/pdf/2602.21877
• Project Page: https://laitifranz.github.io/MemCoach/
• Github: https://laitifranz.github.io/MemCoach/

✨ Datasets citing this paper:
• https://huggingface.co/datasets/laitifranz/MemBench-InternVL3.5-Eval
• https://huggingface.co/datasets/laitifranz/MemBench

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#PhotoMemorability #MLLMs #ComputerVision #AIResearch #HumanComputerInteraction
✨InfoPO: Information-Driven Policy Optimization for User-Centric Agents

📝 Summary:
InfoPO optimizes agent-user collaboration for underspecified requests. It uses an information-gain reward to credit valuable turns that reduce uncertainty, improving decision-making and outperforming multi-turn RL baselines.

🔹 Publication Date: Published on Feb 28

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.00656
• PDF: https://arxiv.org/pdf/2603.00656
• Github: https://github.com/kfq20/InfoPO

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#ReinforcementLearning #AI #HumanComputerInteraction #InformationTheory #AIagents
✨MIBURI: Towards Expressive Interactive Gesture Synthesis

📝 Summary:
MIBURI is an online, real-time framework generating expressive full-body gestures and facial expressions for spoken dialogue. It uses body-part aware codecs and LLM embeddings to create natural, diverse, and contextually aligned motions causally, overcoming limitations of prior methods.

🔹 Publication Date: Published on Mar 3

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.03282
• PDF: https://arxiv.org/pdf/2603.03282
• Project Page: https://vcai.mpi-inf.mpg.de/projects/MIBURI/
• Github: https://github.com/m-hamza-mughal/miburi

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#GestureSynthesis #AI #HumanComputerInteraction #NLP #RealtimeTech
✨ReactMotion: Generating Reactive Listener Motions from Speaker Utterance

📝 Summary:
This paper introduces ReactMotion, a framework for generating natural listener body motions that react appropriately to speaker utterances. It uses a large dataset and preference-based training to create diverse, realistic responses, outperforming prior methods.

🔹 Publication Date: Published on Mar 16

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.15083
• PDF: https://arxiv.org/pdf/2603.15083

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#AI #MachineLearning #HumanComputerInteraction #GenerativeAI #ComputerAnimation
❤1
✨AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents

📝 Summary:
The paper presents AndroTMem, a framework and benchmark diagnosing interaction memory failures in long-horizon GUI agents. It proposes Anchored State Memory ASM, which uses causally linked intermediate-state anchors to overcome this bottleneck, improving task completion rates by up to 30%.

🔹 Publication Date: Published on Mar 19

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.18429
• PDF: https://arxiv.org/pdf/2603.18429

==================================

For more data science resources:
✓ https://xn--r1a.website/DataScienceT

#GUIAgents #AIMemory #AIAgents #AIResearch #HumanComputerInteraction
AI & ML Papers
Photo
🔥 GNM Head: A Generative aNthropometric Model of the human head

💡 The paper introduces a new parametric model called the Generative Anthropometric Model of the human head, or GNM Head. The model is designed to address the limitations of existing publicly available models, which typically only capture the outer geometry of the head and ignore internal structures such as the eyes and mouth. These existing models also often suffer from reduced geometric quality due to low-fidelity input data sets.

The GNM Head model is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy-specific artist-made samples. The model encompasses the head, face, neck, eyeballs, teeth, and tongue, and includes specialized sub-models for the ocular and intra-oral structures.

The paper details the data provenance, model architecture, and performance of the GNM Head model, including its ability to fit target 3D face scans. The results show that the GNM Head model is a significant improvement over existing models, offering a more comprehensive and accurate representation of the human head.

To foster community innovation, the complete GNM framework is made publicly available. The GNM Head model has the potential to serve as a crucial conditioning signal within generative large vision models, allowing for tight spatial control of generated imagery. Overall, the paper contributes a new and improved parametric model of the human head, which can be used in a variety of applications, including computer vision, graphics, and animation.


📅 Published on Jul 26

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23687
• PDF: https://arxiv.org/pdf/2607.23687

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#GenerativeModeling #AnthropometricAnalysis #3DHeadModeling #FacialReconstruction #HumanComputerInteraction
❤2
🔥 MatrAIx: Simulating the World with 8.3 Billion Persona Agents

💡 The paper introduces MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. The problem addressed is that human evaluation of AI systems and digital products is costly, slow, and difficult to scale, while offline evaluations often abstract away human diversity and interactive behavior.

The MatrAIx infrastructure has three core components: Persona8B, which contains 8.3 billion persona records represented by 1290 categorical dimensions, with a quality-filtered core set of approximately 1 million personas comprising 599847 human-grounded and 400000 synthetic records. The MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. MatrAIx also provides 1010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare.

The persona agents were powered by three large language models: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The results of 18189 evaluation trials across eight representative tasks showed that decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance.

The paper also conducted two main validation studies. The first study, a 400-trial controlled study, evaluated persona adherence across ten behavioral attributes and all four environments, with declared behavior expressed or correctly suppressed in 366 trials, which is 91.5 percent. The second study had human and LLM judges evaluate the extraction quality of human-grounded personas.

Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users, addressing the need for scalable and realistic evaluation of AI systems and digital products.


📅 Published on Aug 4

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2608.04205
• PDF: https://arxiv.org/pdf/2608.04205
• Project Page: https://matraix.ai/

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#ArtificialIntelligenceSimulation #PersonaBasedModelling #HumanComputerInteraction #AISystemEvaluation #SimulatedUserTesting
AI & ML Papers
Photo
🔥 MAI-UI Technical Report: Real-World Centric Foundation GUI Agents

💡 The paper addresses the difficulty of moving graphical user interface GUI agents from research prototypes to practical, large‑scale applications. The authors identify four major obstacles: 1 current agents do not allow direct interaction with users, 2 they are limited to operating only within a GUI, 3 there is no established architecture for deploying agents in real environments, and 4 existing agents perform poorly when the number of parallel environments grows.

To overcome these problems the authors introduce a system called MAI‑UI. MAI‑UI is built as a single evolving data pipeline that continuously expands the information used for navigation. The pipeline adds user interaction data and calls to a multi‑component processing module, incorporates a device‑cloud collaboration layer that routes execution based on task state, and includes an online reinforcement‑learning framework that applies advanced optimization to scale both the number of parallel environments and the length of each task. The design is intended to work for a wide range of model sizes, from two billion to two hundred thirty‑five billion parameters, and to support both desktop and mobile navigation.

The experimental results show that MAI‑UI establishes a new state of the art across several standard GUI benchmarks. On the Screen Spot Pro benchmark it achieves a success rate of seventy‑three point five percent, on the MMBench GUI L2 benchmark ninety‑one point three percent, on the OS World G benchmark seventy‑zero point nine percent, and on the UI Vision benchmark forty‑nine point two percent, all of which exceed the previous best results from Gemini‑3‑Pro and Seed‑1‑8. For mobile GUI navigation the system reaches a new state of the art of seventy‑six point seven percent on Android World, surpassing UI‑Tars‑2, Gemini‑2‑5‑Pro and Seed‑1‑8. On the Mobile World benchmark MAI‑UI obtains a success rate of forty‑one point seven percent, outperforming end‑to‑end GUI models and competing favorably with the Gemini‑3‑Pro agent framework.

Additional online reinforcement‑learning experiments demonstrate that scaling the number of parallel environments from thirty‑two to five hundred twelve yields an improvement of five point two in performance, while increasing the environment step budget from fifteen to fifty adds another four point three. The device‑cloud collaboration component improves on‑device performance by thirty‑three percent, cuts the number of cloud model calls by more than forty percent, and preserves user privacy.

Overall the contribution of the paper is a unified methodology that enables large‑scale, privacy‑preserving, and highly effective GUI agents for both desktop and mobile platforms, and it provides empirical evidence that the approach outperforms existing state‑of‑the‑art systems across multiple benchmarks.


📅 Published on Dec 26, 2025

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2512.22047
• PDF: https://arxiv.org/pdf/2512.22047

🤖 Models citing this paper:
• https://huggingface.co/Tongyi-MAI/MAI-UI-8B
• https://huggingface.co/mlx-community/MAI-UI-8B-bf16
• https://huggingface.co/mlx-community/MAI-UI-8B-6bit

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#GUIAgents #HumanComputerInteraction #ScalableAI #AgentArchitecture #RealWorldDeployment
👍1