Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I argue the barrier to FPGA adoption is the actual use-case, not the tools. People who need them can use the tools just fine. FPGA designers are not idiots. They do not walk around scratching their heads wondering, "how come nobody is using our products? If only they were easier to use!" There is a new open-source attempt at replacing imcumbent tools bi-annually begun and dropped.

The use cases for FPGAs are a much harder impediment to adoption. Many people get "FPGA boners" when they even hear the word, fancying themselves "chip designers," but practical use cases are much rarer. As evidence, notice they predominate in the military world, where budget is less of an issue than the commercial world.

The technical issue with FPGAs is that they are still one level abstracted from any CPU. They are only valuable in problems where some algorithm or task can be done with specific logic more quickly than the CPU, given 1. the performance hit of reduced real estate and 2. reduced clock speed relative to a CPU and 3. more money than a CPU.

Further diminishing their value is that any function important enough to require an FPGA can more economically get absorbed into the nearest silicon. For example, consider the serial/deserial coding of audio/video codecs. That used to be done in FGPAs, but got moved into a standard bus (SPI) and moved into codecs and CPUs.

Because of this rarity, experienced engineers know that when an FPGA is introduced to the problem in practical reality, it's a temporary solution (most often to make time-to-market). This confers a degree of honor which is why people get so emotionally-aroused about FPGAs.

You can bet though, that if whatever search function Microsoft is running on those FPGAs proves to be useful, it will be soon absorbed into a more economical form, such as an ASIC, or, more likely, additional instructions to the CPU.

Really, installs on 1,600 servers such as this article reports, is not that impressive and certainly only a prototypical rollout.



Ok, I'll take the counter argument.

FPGAs promise the designer 'arbitrary logic' and deliver 'a place others sell into.'

I disagree that FPGA experts "like" the tools they are given, they tolerate them. One of my friends worked at Xilinx for 15 years and understood this all too well. He felt the leading cause of the problem was that the tools group was a P&L center, they needed to turn a profit in order to exist. They got that profit by charging high prices for the tools and high prices for support. His argument was that 'easier' tools cut into support revenue. When I've had high level (E-level, but not C-level) discussions with Xilinx and Altera there has been a lot of acknowledgement about the 'difficulty of getting up to speed' on the tool chain and many free hours of consulting are offered. From a business engagement point of view, making hard to use tools and then "giving away" thousands of dollars of free consulting to the customer to gain their support seems to work well. The customer feels supported, and stops wondering why if the have consultants around for free those consultants wouldn't just make the tools more straight forward to use and available on a wider variety of platforms.

But the biggest thing has always been intellectual property. You buy an STM32F4 and it has an Ethernet Mac on it (using Synopsis IP as evidenced by the note in the documentation), you pay $8 for the microprocessor, work around the bugs, and get it running. If you buy an FPGA, lets say a Spartan 3E, you pay $18 for the chip, and if you want to use that Synopsis Ethernet MAC?[1] $25,000 for the HDL source to add to our project $10,000 if you are ok with just the EDIF output which can be fed into a place-and-route back end. Oh and some royalty if you ship it on a product you are selling.

The various places that have been accumulating 'open' IP such as Open Cores (http://opencores.org/)have been really helpful for this but it really needs a different pricing model I suspect. A lot of HDL is where OS source was back at the turn of the century (locked down and expensive).

[1] I did this particular exercise in 2005 when I was designing a network attached memory device (https://www.google.com/patents/US20060218362) and was appalled at the extortionate pricing.


> From a business engagement point of view, making hard to use tools and then "giving away" thousands of dollars of free consulting to the customer to gain their support seems to work well.

To me, the entire recent history of computer industry (well, all of it is recent BTW) shows that, if you want your technology to become mass-adopted, you need to make it easier for the little guy to get in the game. The high school kid tinkering with stuff in the parents' basement; the proverbial starving student. That's how x86 crushed RISC; that's how Linux became prominent; that's how Arduino became the most popular micro-con platform (despite more clever things being available).

You make the learning curve nice and gentle, and you draw into your ranks all the unwashed masses out there. In time, out of those ranks the next tech leaders will emerge.


I don't disagree, and I suggested as much to the Xilinx folks (well their EVP of marketing at the time) that if they just added $0.25 to the price per chip they could fund the entire tools effort with that 'tax' and since they would be 'giving away' the tools they could re-task all of the compliance guys who were insuring that licenses worked or didn't work into building useful features.

Their counter is of course that they have customers who sweat the $0.25 difference in price. (which I understand but $10,000 in tools and $15,000 in consulting a year is a hundred thousand chips. Which they say "oh at that volume we would wave the tooling cost." And that got me back to your point of "You already have their design win, why give them free tools? Why not give free tools who have yet to commit to your architecture?"

It is a very frustrating conversation to have.


What's your opinion on the new xilinx c-base design tools, and altera's opencl tools for doing compute acceleration ?


I haven't used either of them. I played around with Systems C a bit when it was the rage but found that my issues weren't in optimizing some bit of C code with a better opcode rather it was assembling a system with the peripherals I wanted in the places I wanted them.

For a long time I considered soft CPUs a bad idea (the Stretch guys kept trying to sell me on them but since I wasn't really doing things like deep packet inspection I didn't have a good use case, even RAID algorithms on them were better handled by pretty generic DSP type architectures.) However in playing with the Zedboard which has a couple of Cortex A9's attached to the Xilinx fabric I find some interesting things there. If only as a new kind of 'i/o' but that is neither I/O port based nor memory map based (it expresses as memory but it feels different than the memory mapping of old like on the PDP/VAX machines and 68K systems). Could just be nostalgia though.


In video industry FPGA based devices are used relatively wide. Multiplexers, encoders, satellite and cable devices etc. And they don't do "everything", usually there is a regular purpose cpu - power, x86, arm etc. as a controller, which also hosts OS and all software and lots of specialized FPGAs. In one device there can be one low power dual core cpu (celeron, old ppc) and tens of top level monster FPGAs like Stratix V.

Reconfiguration is also very much needed - customers often experience network and different specific problems that are only fixable in hardware. So we need configurable packet processors (basically top level router in a chip), configurable processing fpga, lots of small fpga etc. And bug fix for one customer is then gradually distributed to everyone. And devices lifespan is long - many customers use 5 or even 10 year old hardware. Performance is enough for them and new functionality is provided for them, partially in way of new fpga firmware.

PS: some devices, especially cable or satellite, are sold in thousands by several companies, so its not like unique hand-made hardware.

PPS: of course any costs to buy toolchain or windows pc in such companies doesn't matter really. Finding talented fpga designers is way harder as far as I understand.


And, according to the research paper (http://research.microsoft.com/pubs/212001/Catapult_ISCA_2014...) the benefits aren't especially convincing.

- 95% throughput increase at same tail latency, - 29% tail latency improvement at same throughput

I would have expected factors or orders of magnitude.


Thanks. Interesting paper.

Further ,this increases total cost of ownership by 30%, So the performance improvement is just about 70%.


A big advantage that FPGAs have over ASICs, aside from time-to-market and one-time (fab) costs, is reconfigurability. Take a few seconds to stream in a new bitfile and suddenly you have a different chip. This seems like an obvious need in something like a datacenter full of FPGAs: unless you've purpose-built the thing for one or a small set of algorithms that never change, you want the ability to deploy a completely new logic design tomorrow or next week or next year.


What type of logic are you going to be changing every other week in the Datacenter?

What chillingeffect is saying (and I agree) is that this logic should be running on the cpu.


Actual chip design needs FPGAs for hardware emulation prior to fabrication. Yes you'll change the logic out on a weekly, daily, or even hourly basis.


Yes and no. A small ASIC can be prototyped on an FPGA, but FPGA are radically smaller and slower than high end chips.


Well yes, of course, but that's a case where (as mentioned above) the FPGA is a temporary for a final chip.


Except that the MS Catapult results seem to suggest that you can get a substantial timing bump without much power overhead by using a second network of FPGAS running concurrently with the CPUS.

And in the case of Catapult, they'll be refactoring the algorithm to represent changes to their search feature matching, which is the majority of what they offloaded to FPGAs.


Yes, the Bing example seems to give real-world validation of all this.

In general it seems that hardware vs. software fit is orthogonal to concerns such as frequent reconfiguration. Some algorithms simply match hardware well (high concurrency, low-complexity control flow, able to stream through data without complex state). These are often algorithms that do not fit general-purpose CPUs well (cache hierarchy wasted on streaming; lots of control overhead; low core counts relative to FPGA-level parallelism). Some of these algorithms may be for specialized and/or frequently-changing applications such that they should not be burned into an ASIC that will live in a datacenter for 3-5 years.


> any function important enough to require an FPGA can more economically get absorbed into the nearest silicon. For example, consider the serial/deserial coding of audio/video codecs. That used to be done in FGPAs, but got moved into a standard bus (SPI) and moved into codecs and CPUs.

But isn't that because, back then, CPUs simply weren't fast enough for decent video coding? And they still aren't blazing fast for that purpose - compared to GPUs (see GPU-enabled video encoders).

I think the main argument against FPGAs is that it's still a chore to re-purpose them. Sure, it's not as painful as making a new chip from sand, but it's harder than applying code changes to software running in production on CPUs.

If it's a relatively simple algorithm, that changes rarely, where speed is the main bottleneck, that's a very promising scenario for an FPGA.


FPGA in the datacenter may be a red herring; a relatively small percentage of digital circuits are computers.

FPGA is widely used in control circuits for industrial uses, for example.


It is instructive to note where FPGAs are used (and yes, they are used). For example, lower-priced oscilloscopes. Too small volume to merit an ASIC, too high performance to use a CPU, and too fast to interface directly to the central DSP.


FPGAs really only make sense when you need to move a metric assload of data around in a hurry while simultaneously running "embarassingly parallel" logic... and even then, they are rarely the most economical approach in the long run. As a rule they are replaced by ASICs in specialized high-volume applications and CPUs in less specialized ones.

There's nothing better for prototyping and proof-of-concept work, though. And they'll always have a place in low-volume applications that aren't cost-sensitive.


They are also replacing TTL circuits in small volume command and control applications. Circuits with dozens or hundreds of 7400 series ICs wired together in Byzantine ways.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: