I2PS SOLUTIONSENGINEERING BLOGi2psolutions.com

Working Effectively with AI

Part 2 of 3 · Choose the Right AI

A. RemaniFounder, CEO & Head of Engineering · I2PS Engineering

Abstract

AI is not one uniform tool. Different systems are built with different capabilities, interfaces, limits, specializations, and access levels. This field report argues that before deciding an AI is good or bad, the user should first determine whether it is the right system for the task, understand what it can and cannot do, test it on representative work, and revisit old conclusions as the technology changes. Tool selection is part of AI competence.

Keywords: artificial intelligence, model selection, specialized AI, general-purpose AI, capability matching, AI evaluation, subscriptions, human judgment.

Part One of this series argued for AI redundancy: important work can benefit when one model produces and another model independently reviews. But redundancy immediately creates another question. Which models should be in the workflow in the first place?

That question is more important than it looks. AI is often discussed as though every model were a different brand of the same product. In practice, that is not how these systems behave.

Some systems are broad generalists. Some are optimized around coding, research, images, audio, video, automation, or other narrower purposes. Some can use external tools. Some cannot. Some accept particular file types or large working contexts. Some are strongest when the task stays inside a specific domain.

Before saying that an AI failed, the user should first ask whether the right AI was chosen.

This Is Not a Ranking

This article is not a permanent ranking of AI companies or models. A ranking would age quickly, and our own use is too specific to justify declaring one universal winner.

A system that is excellent for one workflow may be a poor choice for another. A model that was limited six months ago may receive a major upgrade. A tool that once led a category may later be overtaken. Access levels can also change what a user is actually testing.

The useful question is therefore not, “Which AI is best?”

It is: “Which AI is best suited to this task, under these constraints, today?”

Ask “Best for What?”

If the objective is to edit video, a system with no meaningful video capability should not be treated as the reference test for AI video work. If the task is software deployment, a model that cannot inspect the relevant files or tools may be at an immediate disadvantage. If the work depends on a narrow technical domain, a specialized system may outperform a broader one even if the broader model is more capable overall.

That does not make the general-purpose model bad. It makes the match poor.

Engineers already think this way about physical tools. A multimeter, an oscilloscope, a torque wrench, and a thermal camera can all be excellent tools. None becomes defective because it cannot replace the others.

AI should be evaluated with the same discipline.

Capability Mismatch Is Not AI Failure

A capability mismatch can produce a misleading conclusion. The user asks the system to distinguish, generate, inspect, or manipulate something it was not designed or equipped to handle well, then treats the weak result as proof that the entire technology is weak.

“It’s like asking a colorblind person to choose from a palette of colors and name each one of them.”

— A. Remani

The analogy is deliberately simple. If the task depends on a capability the evaluator does not possess, more insistence does not create the missing capability. In the same way, repeatedly prompting the wrong AI may only produce more elaborate versions of the same mismatch.

The first responsibility of the user is therefore selection.

Generalists and Specialists

General-purpose AI has enormous value because one system can help across writing, analysis, programming, planning, translation, research, and many everyday tasks. For a small company, that breadth can be economically powerful.

But breadth has limits. A specialized tool may have interfaces, training, workflows, integrations, or output controls that a general model does not provide. The best workflow may therefore combine a capable generalist with one or more specialists rather than force one model to imitate an entire software stack.

This is not redundancy for its own sake. It is division of labor.

The user should decide whether the task needs a generalist, a specialist, or a combination.

Understand the AI Before You Judge It

Choosing a model is only the beginning. The user also needs to understand the system being tested.

What kinds of files can it actually inspect? Which tools are available in the current plan? Can it browse, execute code, work with images, generate media, maintain a long context, or connect to external services? Are those abilities available in the interface being used, or only in another product tier or mode?

A user does not need to memorize a technical manual. But testing a system without understanding its basic capabilities can make the test meaningless.

Sometimes what appears to be a model limitation is really an interface limitation. Sometimes it is a subscription limitation. Sometimes the feature exists but has not been enabled in the workflow. And sometimes the model truly is the wrong tool.

Those are different conclusions and should not be confused.

Free Access, Paid Access, and Fair Testing

Cost matters, especially for startups and individual users. But a fair evaluation does not require buying every premium plan on the market.

Where a useful free tier or trial exists, start there. Give the system representative work rather than artificial demonstrations. Learn how it responds. Identify which limitations actually affect the task.

If a paid level unlocks a capability that materially matters, it can be reasonable to pay for the minimum appropriate access long enough to evaluate it properly. The goal is not to collect subscriptions. The goal is to avoid rejecting a tool because the version tested was never capable of performing the intended workflow.

The same principle works in the opposite direction. A premium subscription is not evidence that a system is the right choice. If a cheaper or free tool performs the required task reliably, paying more does not make the workflow more intelligent.

Value comes from fit.

Test the Work You Actually Need

Public benchmarks can be useful, but they do not know your exact job.

A startup working on multilingual websites, technical documents, engineering prototypes, software changes, and field research has a different evaluation problem from a video studio, a legal team, a classroom, or a data-science group.

The most useful test is therefore representative work.

Give the model a real but controlled task. Measure whether it follows constraints. Check the delivered artifact rather than the explanation. Record how much correction was required. Compare speed, reliability, usability, and the amount of human supervision needed.

A model that produces an impressive first answer but requires hours of repair may have less practical value than one whose first answer looks less dramatic but survives verification.

Capabilities Change, So Old Conclusions Expire

AI changes quickly enough that users should be careful with permanent judgments.

“I tried that AI and it was bad” is incomplete without a date, a model, a plan, a mode, and a task.

The system may have changed. The feature set may have changed. The model behind the product may have changed. A capability that did not exist during the original test may now be normal.

That does not mean users should continuously chase every release. It means important tool decisions should be revisited when the underlying technology materially changes.

The best AI for a task is a moving target.

One AI Does Not Need to Do Everything

There is a strange expectation around AI that does not exist with most other technology: because a model can do many things, users expect it to do everything.

That expectation can create unnecessary disappointment and unnecessary risk.

A general model can help define a video concept while a video-focused system produces the media. One model can write code while another audits the deployment. A research-capable system can gather evidence while a different tool handles structured analysis. Human specialists can enter wherever domain judgment is required.

The workflow should be designed around the objective, not around proving that one subscription can perform every role.

Avoid the Sameness Trap

There is another reason tool selection matters. If everyone chooses the same general-purpose system, gives it the same vague instruction, and accepts the first answer, AI can create a surprising amount of sameness.

The same structures appear. The same phrasing appears. The same obvious ideas appear. The technology becomes an equalizer, but it can also flatten differences between users.

That is not inevitable.

Two people with access to the same AI can produce very different work because one understands the domain, chooses a more appropriate tool, combines systems intelligently, provides better evidence, rejects generic output, and applies independent judgment.

Access to AI may become common. Knowing how to assemble the right AI workflow is a separate skill.

Selection Is a Human Responsibility

The AI cannot be fully responsible for whether it was the right AI to choose. That decision sits one level above the model.

The user defines the objective, evaluates the available tools, decides what capabilities matter, accepts the cost, and determines whether the system performed well enough.

That is why AI literacy should include more than prompting. It should include tool selection, capability awareness, testing, cost awareness, and the willingness to change tools when the work requires it.

Part Three: The User Might Be the Problem

Choosing the right AI still does not guarantee a good result.

A capable model can be given a vague objective. A strong coding system can be asked to change too much at once. A powerful research model can be fed poor assumptions. A correct first result can be destroyed by a long chain of contradictory corrections.

That leads to the final part of this series.

Part One asks whether important AI work should have independent redundancy. Part Two asks whether the right AI was selected. Part Three asks the uncomfortable question that remains after both of those conditions are satisfied:

What if the user is the problem?

About the Author

A. Remani

A. Remani is the Founder, CEO, and Head of Engineering at I2PS Engineering. His work combines engineering, product development, software, field operations, prototyping, and the practical integration of AI into small-team workflows. His writing focuses on responsible engineering, system reliability, and practical knowledge transfer.

Closing Reflection

Before judging an AI, make sure you are judging the right tool, in the right mode, on the right task. Tool selection is not separate from AI competence; it is part of it.