The Department of Labor released an AI Literacy Framework today, and it's really encouraging to see that the approach aligns closely with the work we're doing at AI for Education. The framework defines AI literacy with these foundational content areas: - Understanding AI principles - Exploring AI uses - Directing AI effectively - Evaluating outputs - Using AI responsibly What's notable is that the framework addresses both content and delivery, which we think about all the time. Beyond the foundational content areas, it emphasizes effective delivery principles like experiential learning, building agility, and developing complementary human skills. We also like that the framework doesn't focus primarily on risks or displacement. Instead, it positions AI as a tool that can support workers and open new opportunities when people have the right skills and support. It acknowledges real concerns about AI's workforce impact while focusing on actionable preparation. It's encouraging to see the DOL support the kind of practical, human-centered AI literacy work that we have seen work with our partners. You can find the link to the full framework in the comments. Let us know what you think. #AILiteracy #FutureOfWork #AIinEducation
AI Frameworks For Software Development
Explore top LinkedIn content from expert professionals.
-
-
McKinsey's report on 'AI Bank of the Future' - with a solid updated. What’s new? 🔹𝐌𝐨𝐫𝐞 𝐜𝐨𝐦𝐩𝐥𝐞𝐭𝐞 𝐨𝐧 𝐆𝐞𝐧𝐀𝐈 𝐚𝐫𝐜𝐡𝐢𝐭𝐞𝐜𝐭𝐮𝐫𝐞 From tools ➜ to agents From analytics ➜ to orchestration From point solutions ➜ to full-stack, GenAI-native architecture At the center of it all is 𝐦𝐮𝐥𝐭𝐢𝐚𝐠𝐞𝐧𝐭 𝐨𝐫𝐜𝐡𝐞𝐬𝐭𝐫𝐚𝐭𝐢𝐨𝐧, where specialized AI agents are coordinated by a central AI “brain” that can plan, route tasks, and trigger the right agent at the right time. 🔹 𝐅𝐫𝐨𝐦 “𝐮𝐬𝐞 𝐜𝐚𝐬𝐞𝐬” 𝐭𝐨 𝐝𝐨𝐦𝐚𝐢𝐧 𝐫𝐞𝐰𝐢𝐫𝐢𝐧𝐠 Instead of building scattered pilots, the report suggests identifying a few high-impact subdomains (like credit, onboarding, or fraud) and transforming them end to end. 🔹𝐒𝐭𝐫𝐨𝐧𝐠𝐞𝐫 𝐞𝐦𝐩𝐡𝐚𝐬𝐢𝐬 𝐨𝐧 𝐫𝐞𝐮𝐬𝐞 𝐚𝐧𝐝 𝐬𝐜𝐚𝐥𝐞 Reusable AI components. Shared orchestration logic. AI control towers to drive governance and coordination across teams. In short: the vision is bigger, but the path is more detailed. It’s still early for most banks (and probably also most industries) to adopt a stack like this. But the ideas here: agents, orchestration, domain-level transformation are starting to feel more widely relevant. We’ve been thinking about similar challenges on our end too. 📍We recently open-sourced an internal project we’ve been building, 𝐆𝐞𝐧𝐀𝐈 𝐀𝐠𝐞𝐧𝐭𝐎𝐒 - a lightweight framework for orchestrating multiagent systems. If you’re exploring in this direction, check it out: GitHub https://bit.ly/4kzE1Mt And if you're also a fan of open source, giving it a ⭐ would mean a lot to us! PS. Curious how many teams are thinking in this direction too. __________ For more on AI and open-source materials, plz check my previous posts. I share my journey here. Join me and let's grow together. Alex Wang #agenticai #aiagents #artificialintelligence #technology
-
Agentic AI is 𝗻𝗼𝘁 about wrapping prompts around a large language model. It’s about designing systems that can: → 𝗣𝗲𝗿𝗰𝗲𝗶𝘃𝗲 their environment → 𝗣𝗹𝗮𝗻 actionable steps → 𝗔𝗰𝘁 on those plans → 𝗟𝗲𝗮𝗿𝗻 and improve over time And yet, many teams hit a wall—not because the models fail, but because the 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 behind them isn’t built for agent behavior. If you’re building agents, you need to think in 𝗳𝗼𝘂𝗿 𝗱𝗶𝗺𝗲𝗻𝘀𝗶𝗼𝗻𝘀: 1. 𝗔𝘂𝘁𝗼𝗻𝗼𝗺𝘆 & 𝗣𝗹𝗮𝗻𝗻𝗶𝗻𝗴 → Agents must decompose goals into steps and execute them independently. 2. 𝗠𝗲𝗺𝗼𝗿𝘆 & 𝗖𝗼𝗻𝘁𝗲𝘅𝘁 → Without memory, agents forget past context. Vector DBs like FAISS, Redis, or pgvector aren’t optional—they’re foundational. 3. 𝗧𝗼𝗼𝗹 𝗨𝘀𝗮𝗴𝗲 & 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻 → Agents must go beyond text generation—calling APIs, browsing, writing code, and executing it. 4. 𝗖𝗼𝗼𝗿𝗱𝗶𝗻𝗮𝘁𝗶𝗼𝗻 & 𝗖𝗼𝗹𝗹𝗮𝗯𝗼𝗿𝗮𝘁𝗶𝗼𝗻 → The future isn’t just one agent. It's many, working together—planner-executor setups, sub-agents, role-based dynamics. Frameworks like 𝗟𝗮𝗻𝗴𝗚𝗿𝗮𝗽𝗵, 𝗔𝘂𝘁𝗼𝗚𝗲𝗻, 𝗟𝗮𝗻𝗴𝗖𝗵𝗮𝗶𝗻,𝗚𝗼𝗼𝗴𝗹𝗲'𝘀 𝗔𝗗𝗞, and 𝗖𝗿𝗲𝘄𝗔𝗜 make these architectures more accessible. But frameworks alone aren’t enough. If you’re not thinking about: • 𝗧𝗮𝘀𝗸 𝗱𝗲𝗰𝗼𝗺𝗽𝗼𝘀𝗶𝘁𝗶𝗼𝗻 • 𝗦𝘁𝗮𝘁𝗲𝗳𝘂𝗹𝗻𝗲𝘀𝘀 • 𝗥𝗲𝗳𝗹𝗲𝗰𝘁𝗶𝗼𝗻 • 𝗙𝗲𝗲𝗱𝗯𝗮𝗰𝗸 𝗹𝗼𝗼𝗽𝘀 …your agents will likely remain shallow, brittle, and fail to scale. The future of GenAI lies in 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝗶𝗻𝗴 𝗶𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝘁 𝗯𝗲𝗵𝗮𝘃𝗶𝗼𝗿, not just fine-tuning prompts. 2025 is the year we go from 𝗽𝗿𝗼𝗺𝗽𝘁 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 to 𝗔𝗜 𝘀𝘆𝘀𝘁𝗲𝗺 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘀. Let’s build agents that don’t just respond—but 𝗿𝗲𝗮𝘀𝗼𝗻, 𝗮𝗱𝗮𝗽𝘁, 𝗮𝗻𝗱 𝗲𝘃𝗼𝗹𝘃𝗲.
-
I’ve noticed that many GenAI application projects put in automated evaluations (evals) of the system’s output probably later — and rely on humans to manually examine and judge outputs longer — than they should. This is because building evals is viewed as a massive investment (say, creating 100 or 1,000 examples, and designing and validating metrics) and there’s never a convenient moment to put in that up-front cost. Instead, I encourage teams to think of building evals as an iterative process. It’s okay to start with a quick-and-dirty implementation (say, 5 examples with unoptimized metrics) and then iterate and improve over time. This allows you to gradually shift the burden of evaluations away from humans and toward automated evals. I wrote previously in The Batch about the importance and difficulty of creating evals. Say you’re building a customer-service chatbot that responds to users in free text. There’s no single right answer, so many teams end up having humans pore over dozens of example outputs with every update to judge if it improved the system. While techniques like LLM-as-judge are helpful, the details of getting this to work well (such as what prompt to use, what context to give the judge, and so on) are finicky to get right. All this contributes to the impression that building evals requires a large up-front investment, and thus on any given day, a team can make more progress by relying on human judges than figuring out how to build automated evals. I encourage you to approach building evals differently. It’s okay to build quick evals that are only partial, incomplete, and noisy measures of the system’s performance, and to iteratively improve them. They can be a complement to, rather than replacement for, manual evaluations. Over time, you can gradually tune the evaluation methodology to close the gap between the evals’ output and human judgments. For example: - It’s okay to start with very few examples in the eval set, say 5, and gradually add to them over time — or subtract them if you find that some examples are too easy or too hard, and not useful for distinguishing between the performance of different versions of your system. - It’s okay to start with evals that measure only a subset of the dimensions of performance you care about, or measure narrow cues that you believe are correlated with, but don’t fully capture, system performance. For example if, at a certain moment in the conversation, your customer-support agent is supposed to (i) call an API to issue a refund and (ii) generate an appropriate message to the user, you might start off measuring only whether or not it calls the API correctly and not worry about the message. Or if, at a certain moment, your chatbot should recommend a specific product, a basic eval could measure whether or not the chatbot mentions that product without worrying about what it says about it. [Truncated due to length limit. Full text: https://lnkd.in/gygj3y7w ]
-
McKinsey & Company 𝗮𝗻𝗮𝗹𝘆𝘇𝗲𝗱 𝟭𝟱𝟬+ 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗚𝗲𝗻𝗔𝗜 𝗱𝗲𝗽𝗹𝗼𝘆𝗺𝗲𝗻𝘁𝘀 — 𝗮𝗻𝗱 𝗳𝗼𝘂𝗻𝗱 𝗼𝗻𝗲 𝗰𝗼𝗺𝗺𝗼𝗻 𝘁𝗵𝗿𝗲𝗮𝗱: ⬇️ One-off solutions don’t scale. The most successful projects take a different path: They use open, modular architectures that enable speed, reuse, and control. → Designed for reuse → Able to plug in best-in-class capabilities → Free from vendor lock-in This is the reference architecture McKinsey now recommends — optimized to scale what works while staying compliant. It consists of five core components: ⬇️ 𝟭. 𝗦𝗲𝗹𝗳-𝘀𝗲𝗿𝘃𝗶𝗰𝗲 𝗽𝗼𝗿𝘁𝗮𝗹: → A secure, compliant “pane of glass” where teams can launch, monitor, and manage GenAI apps. → Preapproved patterns, validated capabilities, shared libraries. → Observability and cost controls built-in. 𝟮. 𝗢𝗽𝗲𝗻 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 → Services are modular, reusable, and provider-agnostic. → Core functions like RAG, chunking, or prompt routing are shared across apps. → Infra and policy as code, built to evolve fast. 𝟯. 𝗔𝘂𝘁𝗼𝗺𝗮𝘁𝗲𝗱 𝗴𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲 𝗴𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀 → Every prompt and response is logged, audited, and cost-attributed. → Hallucination detection, PII filters, bias audits — enforced by default. → LLMs accessed only through a centralized AI gateway. 4. 𝗙𝘂𝗹𝗹-𝘀𝘁𝗮𝗰𝗸 𝗼𝗯𝘀𝗲𝗿𝘃𝗮𝗯𝗶𝗹𝗶𝘁𝘆 → Centralized logging, analytics, and monitoring across all solutions → Built-in lifecycle governance, FinOps, and Responsible AI enforcement → Secure onboarding of use cases and private data controls → Enables policy adherence across infrastructure, models, and apps 5. 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗴𝗿𝗮𝗱𝗲 𝗨𝘀𝗲 𝗖𝗮𝘀𝗲𝘀 → Modular setup for user interface, business logic, and orchestration → Integrated agents, prompt engineering, and model APIs → Guardrails, feedback systems, and observability built into the solution → Delivered through the AI Gateway for consistent compliance and scale The message is clear: If your GenAI program is stuck, don’t look at the LLM. Look at your platform. 𝗜 𝗲𝘅𝗽𝗹𝗼𝗿𝗲 𝘁𝗵𝗲𝘀𝗲 𝗱𝗲𝘃𝗲𝗹𝗼𝗽𝗺𝗲𝗻𝘁𝘀 — 𝗮𝗻𝗱 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝘆 𝗺𝗲𝗮𝗻 𝗳𝗼𝗿 𝗿𝗲𝗮𝗹-𝘄𝗼𝗿𝗹𝗱 𝘂𝘀𝗲 𝗰𝗮𝘀𝗲𝘀 — 𝗶𝗻 𝗺𝘆 𝘄𝗲𝗲𝗸𝗹𝘆 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿. 𝗬𝗼𝘂 𝗰𝗮𝗻 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲 𝗵𝗲𝗿𝗲 𝗳𝗼𝗿 𝗳𝗿𝗲𝗲: https://lnkd.in/dbf74Y9E
-
If you're getting into LLMs, PyTorch is essential. And lot of folks asked for beginner-friendly material over the months, so I put this together: "PyTorch in One Hour: From Tensors to Multi-GPU Training" (https://lnkd.in/ghv-kfkK) 📖 ~1h to read through 💻 Maybe the perfect weekend project!? I’ve spent nearly a decade using, building with, and teaching PyTorch. And in this tutorial, I try to distill what I believe are the most essential concepts. Everything you need to know to get started, and but nothing more, since your time is valuable, and you want to get to building things!
-
I taught myself machine learning > 10 years ago. If I had to start again today, I wouldn’t touch models, LLMs, or agents first, as many AI experts suggest. I'd start with the math and the code. Ugly truth: 90% of people skip the foundations, then wonder why everything feels like magic or falls apart in production. If you want to be different, actually understand ML, not just copy-paste, this is the roadmap I'd follow: Start with fundamentals: Because no matter how fast LLMs or GenAI evolve, your math, code, and logic will keep you relevant. Here's what you should focus on: 📐 1. Linear Algebra Learn these core ideas: Vectors, matrices, tensors Matrix multiplication (dot products, broadcasting) Transpose, inverse, rank, determinants Eigenvalues & eigenvectors (especially for PCA & embeddings) Projections and orthogonality ✅ Use NumPy to implement everything yourself → Practice matrix ops, dot products, and visualizing transformations with Matplotlib 🔁 2. Calculus Focus on: Derivatives & partial derivatives Chain rule (for backpropagation in neural nets) Gradient descent Convex functions, minima/maxima ✅ Use SymPy or JAX to visualize and compute derivatives → Plot functions and their gradients to develop deep intuition 🎲 3. Probability You need a solid grip on: Random variables (discrete & continuous) Conditional probability & Bayes' rule Joint & marginal probability The Chain rule Expectation, variance, entropy Common distributions: Bernoulli, Binomial, Gaussian, Poisson Central limit theorem The law of large numbers ✅ Simulate simple probability experiments in Python with NumPy → E.g. simulate sampling from distributions 📊 4. Statistics These are must-know topics: Descriptive stats: mean, median, mode, standard deviation Hypothesis testing: p-values, confidence intervals, t-tests Correlation vs. causation Sampling, bias, and variance Overfitting/underfitting A/B testing basics ✅ Use Pandas & SciPy to explore real datasets → Calculate descriptive stats, create histograms/box plots, run t-tests 🔧 Essential Python libraries to learn early NumPy – for vectorized math and fast array ops Pandas – for loading, cleaning, and analyzing tabular data Matplotlib / Seaborn – for plotting and visualizing distributions, relationships, and trends SymPy – for symbolic math and calculus SciPy – for stats, optimization, and numerical methods Use Jupyter Notebooks(to combine math, code, & visuals in one place) 📚 Best resources to nail the fundamentals: ✅ Machine Learning Foundations Math series (ML Foundations: Linear Algebra, Calculus, Probability, and Statistics)-series of 4 courses that I've created together with LinkedIn learning ✅ Hands-On ML with TensorFlow & Keras book by Aurélien Géron ✅ The Hundred-page Machine Learning Book by Andriy Burkov If you want to become an actual ML engineer, not just someone who watches and copies demos, start here. ♻️ Repost to help others💚
-
I think Red Hat’s launch of 𝗹𝗹𝗺-𝗱 could mark a turning point in 𝗘𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗔𝗜. While much of the recent focus has been on training LLMs, the real challenge is scaling inference, the process of delivering AI outputs quickly and reliably in production. This is where AI meets the real world, and it's where cost, latency, and complexity become serious barriers. 𝗜𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗶𝘀 𝘁𝗵𝗲 𝗡𝗲𝘄 𝗙𝗿𝗼𝗻𝘁𝗶𝗲𝗿 Training models gets the headlines, but inference is where AI actually delivers value: through apps, tools, and automated workflows. According to Gartner, over 80% of AI hardware will be dedicated to inference by 2028. That’s because running these models in production is the real bottleneck. Centralized infrastructure can’t keep up. Latency gets worse. Costs rise. Enterprises need a better way. 𝗪𝗵𝗮𝘁 𝗹𝗹𝗺-𝗱 𝗦𝗼𝗹𝘃𝗲𝘀 Red Hat’s llm-d is an open source project for distributed inference. It brings together: 1. Kubernetes-native orchestration for easy deployment 2. vLLM, the top open source inference server 3. Smart memory management to reduce GPU load 4. Flexible support for all major accelerators (NVIDIA, AMD, Intel, TPUs) AI-aware request routing for lower latency All of this runs in a system that supports any model, on any cloud, using the tools enterprises already trust. 𝗢𝗽𝘁𝗶𝗼𝗻𝗮𝗹𝗶𝘁𝘆 𝗠𝗮𝘁𝘁𝗲𝗿𝘀 The AI space is moving fast. New models, chips, and serving strategies are emerging constantly. Locking into one vendor or architecture too early is risky. llm-d gives teams the flexibility to switch tools, test new tech, and scale efficiently without rearchitecting everything. 𝗢𝗽𝗲𝗻 𝗦𝗼𝘂𝗿𝗰𝗲 𝗮𝘁 𝘁𝗵𝗲 𝗖𝗼𝗿𝗲 What makes llm-d powerful isn’t just the tech, it’s the ecosystem. Forged in collaboration with founding contributors CoreWeave, Google Cloud, IBM Research and NVIDIA and joined by industry leaders AMD, Cisco, Hugging Face, Intel, Lambda and Mistral AI and university supporters at the University of California, Berkeley, and the University of Chicago, the project aims to make production generative AI as omnipresent as Linux. 𝗪𝗵𝘆 𝗜𝘁 𝗠𝗮𝘁𝘁𝗲𝗿𝘀 For enterprises investing in AI, llm-d is the missing link. It offers a path to scalable, cost-efficient, production-grade inference. It integrates with existing infrastructure. It keeps options open. And it’s backed by a strong, growing community. Training was step one. Inference is where it gets real. And llm-d is how companies can deliver AI at scale: fast, open, and ready for what’s next.
-
You can instantly 10x your AI-generated frontends just by learning what different UI components are called. The Component Gallery (https://component.gallery) is a comprehensive visual reference for standard UI patterns and their proper names ⚡ When delegating tasks to coding agents, the quality of your output is bottlenecked by the specificity of your input. If the only vocabulary you use in your prompts are generic words like "menu" or "button," the model will naturally return generic results. To get high-fidelity interfaces, you need the right shared language for spec-driven development. Knowing when to ask for a "combobox" instead of a "dropdown" or a "segmented control" instead of a "toggle" gives the AI the exact context it needs to execute your vision. The Component Gallery walks through this really well. UI Design Brain is also an open-source taxonomy for UI skills and component design. (https://lnkd.in/gYYbhv5q) based on the Component Gallery. Mastering this terminology is a small investment that massively upgrades your prompt engineering. #ai #programming #softwareengineering
-
I tried EVERY major AI Coding tool so you don’t have to. Here’s what I learned about each one - and which one’s the best for your particular use case 👇 After an entire weekend of hands-on testing 15+ AI coding assistants, building the same real-life application (tax comparison calculator), and documenting every step - here's the comprehensive breakdown to separate the signal from the noise: 🏆 Best Overall: Cline - 100% open source and free version of Cursor + Windsurf that’s a simple VS Code extension - Truly thoughtful agentic coding with extensive tool use (terminal, computer use, websites, etc) - Wrote the best code with fewer mistakes, better self-healing, but no inline chat 🎨 Best for Non-Technical Users: Vercel V0 - Fast, Easy, intuitive UX - Strong community and templates - Component-specific editing via AI is magical ⚡Best for Quick Prototypes: Anthropic Claude 3.5 Sonnet - Fast & clean responses - Great reasoning & logic clarity - Artifact is great for prototyping, with ability to publish and share Replit: Good for full-stack cloud development, but sits in an awkward spot—too complex for beginners, too constrained for advanced users. StackBlitz Bolt.new: A standard cloud IDE with AI codegen, but nothing special. Lovable: Similar to Bolt, but unreliable AI-generated code, hard to toggle/see code. Cursor: Great Copilot alternative, but lacks extensive agentic capabilities like Cline. Codeium Windsurf: Strong agent mode but agent was sometimes lazy and incomplete. GitHub Copilot: Good for simple inline edits, but lacks full agentic workflow (though an agent mode was recently released). Aider: Terminal & keyboard only. Feels like Vim/Emacs on steroids. Too hardcore. OpenHands: Open-source and free Cognition Devin with strong agentic coding, but SaaS version is unstable. OpenAI (o3-mini-high): Good logic depth but lacks a coding canvas. Anthropic (Claude 3.5 Sonnet): Fast + clean. Artifact is great for prototypes, but can’t edit code directly inside it. Google Gemini 2: Poor experience—lazy, incomplete code. Generated separate files that I had to manually combine. DeepSeek AI R1: Strong long reasoning chains, but gets a lot of logic wrong. Tempo (YC S23): Promising PRD → Design → Code → Deploy workflow, but still in early stages. Onlook: Strong for design-first workflows but inconvenient for direct code editing. Reweb: Generates only UI components, not code with logic. My Final Recommendations: - For non-technical users: Vercel V0 is the best no-code/low-code option. - For cloud-based development: Try Bolt. - For local AI-powered coding: Cline is free and outperforms Cursor/Codeium. - For rapid prototyping: Claude 3.5 Sonnet is fast and effective. - For designers: Tempo or Onlook provide a strong UI-first workflow. Do you want to see a full write up of my AI coding experiences? Let me know if I should make a full post comparing AI Coding tools in detail by sharing this post and commenting below.