NEW - ISO 27001 Certification and ONNX and Triton Blueprints for Accelerated Inference
The hub for everything AI, data science, and real-world experimentation.
Recently active
In January I had the pleasure of working with Anbu Valluvan, an HP AI Studio tester, to resolve a GPU limitation issue his team was running into on their personal laptops when Google Colab GPU was not available due to budget constraints or GPU queuing. The team had collected, created data strategies to ingest up and process blood pressure datasets in appropriate sizes to train small models for use with edge devices. However, their access to GPU was limited. To achieve the accuracy requirements for the paper to qualify for peer review and acceptance, Anbu reached out to collaborate on the project using AI Studio because I had access to a NVIDIA GeForce RTX 4070 GPU which allowed him to experiment and prepare the project with multiple iterations when access to NVIDIA A100 GPU was limited.In a couple of minutes Anbu invited me to his AI Studio project, I synced assets and artifacts, and trained the model on my GPU. Because AI Studio saved artifacts, results, and cloned code to Github, Anb
I am looking to run some testing on a local ollama model in AI Studio. However, every time I open AI Studio, I have to redownload ollama and pull llama3 to use from my notebook. Any suggestion for how to save that download (done from terminal) so that it is already loaded when I open AI Studio?
AI is evolving fast, and with all the hype, it’s easy to lose sight of what really matters. In a recent conversation, Thomas H. Davenport and Randy Bean broke down five key AI trends that will shape the year ahead—helping leaders cut through the noise and focus on what’s actually changing the game.n this discussion, they cover:Agentic AI: Hype vs. Reality – How to navigate the excitement and actual impact. AI ROI Matters – Why businesses need to track real value, not just trends. Data-Driven Culture Struggles – The persistent roadblocks to making data central to decision-making. Unstructured Data Challenges – Why tackling messy data is more important than ever. Evolving AI Leadership – How roles and reporting structures are shifting in the AI era.Drawing from their latest MIT Sloan Management Review article, these experts dive into how agentic AI and large language models are reshaping business in 2025. They explore how data-driven decision-making is changing the way companies operate
As a Z by HP Ambassadors, Benedict Neo was equipped with the HP Z6 G5 Tower Workstation, featuring an AMD Ryzen Threadripper PRO 7965WX processor and NVIDIA RTX 6000 graphics. This powerful setup is tailored to significantly enhance his data science projects. However, a new workstation means he needed to establish a fresh data science environment.Here’s a walkthrough into how he configured his workstation, along with valuable tips to help you optimize your own setup for maximum efficiency. In this setup, Benedict walks you through installing PowerShell 7, WSL, Ubuntu, Git, setting up CUDA, and installing Python with Mamba.How to Setup Windows for Data Science HP Z6 G5 SpecsProcessor: AMD Ryzen Threadripper PRO 7965WX 24-Cores 4.20 GHz Installed RAM: 256 GB (255 GB usable) System type: 64-bit operating system, x64-based processor Graphics: NVIDIA RTX 6000 Ada 48 GB 4DP Graphics
This week Anbu Valluvan Devadasan et all published a paper titled “E-SMOTE: Entropy Based Minority Oversampling for Heart Failure and AIDS Clinical Trails Analysis.”I reached out to Anbu to chat about the Entropy-based Synthetic Minority Oversampling Technique (SMOTE) which uses a subset of the dataset from the minority class to address overbias of a majority class. Anbu explained that the team used entropy as a metric to assess influence of a minority class on outcomes and outperform conventional SMOTE oversampling techniques in some cases. Results illustrated that both SMOTE and E-SMOTE techniques should be considered for addressing imbalanced data samples as they can be used to address class imbalance. Anbu is excited and looking forward to dig deeper into which data characteristics affect metrics and how.Anbu’s open to chat about his work. If you’re interested, ping him - linkedin.com/in/anbuvalluvan. S. Veerla, A. V. Devadasan, M. Masum, M. Chowdhury and H. Shahriar, "E-SMOTE: Ent
Maikel Ronnau et al will publish Automatic segmentation and classification of Papanicolaou-stained cells and dataset for oral cancer detection in next month’s Computers in Biology and Medicine.Their research uses a CNN U-Net architecture with ResNet encoder to achieve expert level performance by segmenting and classifying a trained on UFSC OCPap dataset of oral cytology images. The team uses transfer learning for image segmentation to train the model which is used to evaluate more than 1500 images from 52 patients labeled by specialists resulting in an average Dice score of 0.66 and Intersection over Union (IoU) of 0.65. “The CNN model architecture for encoding layers are based on DenseNet-169 while the decoding layers are LinkNet but to replace the regular softmax layer with a temperature scaling softmax with a temperature parameter value of 0.1 to increase the confidence of the predictions and avoid the bias towards the prediction of background pixels. The model’s prediction is furth
GPU availability and job interruption can be an issue for me, so this paper caught my attention. The paper is titled “Mirage: Towards Low-interruption Services on Batch GPU Clusters with Reinforcement Learning” and it proposes a method for reducing service interruptions on GPU clusters for deep learning jobs.The tool is called Mirage. Mirage is a Slurm-compatible foundational model that uses statistical and reinforcement learning (RL) techniques to proactively provision resources to mitigate interruptions. To accomplish this, researchers alter the decoder transformer architecture to create a Mixture of Experts (MoE) of neural networks. The neural networks share the same network architecture to support scaling up of model capacity and determine ‘best-fit experts’ that optimize job trace analysis outputs for resource allocations. What’s really cool about this paper is how it investigates the use of Deep Q-Learning RL techniques to reenforce rewards to improve GPU queue wait times for a f
NOVA on PBS has a special titled “Secrets in Your Data” where they talk about privacy, access, use, and innovations. A fun show to watch for fellow data lovers… and if you’re also a fan of the HBO show “Silicon Valley” like I am, you’ll enjoy the discussion around a decentralized internet… will it take off?NOVA - Secrets In Your Data
Today, I had the pleasure of interviewing Andrew Kemp. Andrew is the head of product marketing for data science and AI at Z by HP. Andrew recruited me into the Z by HP global ambassador program and is responsible for its creation. In this episode, we learn about his experience teaching in North Korea for 4 years as well as how the landscape of hardware and community is changing.
Is there ever a scenario you need to keep data out of the cloud? For me it’s company data we’re looking to retain as IP. I’m curious to hear what types of data others consider sensitive.
We did a quick survey to more than 200 Data Scientists and ITDMs. Sharing some interesting responses about tools used to create AI.
Leaving this joke here and hopes that it brightens your day:There are two types of Data Scientists: 1. Those that can extrapolate from incomplete data
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.