Hacker Newsnew | past | comments | ask | show | jobs | submit | typs's commentslogin

"Comfortably" varies by person. I know many people living fine on less in San Francisco.


Yes, I absolutely know some. At the same time, many weren’t out of financial necessity but maybe financial convenience. I know a few who had a remote job or worked in their hometown even though they had savings and weren’t spending a lot.

They then would move out to a bigger more expensive city and feel more financially comfortable than they would have. This was in first generation immigrant families, which as the article notes, is where this practice is more common.


Some interesting ideas in this, but I think his argument is somewhat undermined by the fact that his main example of computer use has actually gotten much better recently because of RLVR


Perhaps. RL env companies based in the U.S. sell to Chinese labs quite a bit too though (though on a discount, once they're no longer on the frontier)! And it would make sense that a lot of these problems which are based on work in the U.S. enterprise economy would be coming from the U.S.


What did you do around cross-harness testing? I don't see anything in the blog post about what harnesses were used in evaluation. SOTA benchmarks have consistently shown that frontier model performance is quite sensitive to what tools are exposed (e.g. str_replace vs. apply_patch) as the labs are RLing on their own harnesses. Did you do testing of the models in a standard setup or in their native harnesses?


yes well aware :) numbers shown are on "house" harnesses eg codex with gpt and claude code with opus.

fwiw we have examples of each model doing better on NON-house harnesses too - speaking jsut for myself i think the "the labs are RLing on their own harnesses" narrative is kinda overstated if you think through wanting to have any meaningful api business (often eg the labs will give guidance on what is prefered and the agent labs can easily match tool contract to that, which is to say, the "home turf advantage" isnt as large as you think it is if you try a little bit)


What "non-house" harnesses have you found to work best?


What is the "house" harness for minimax? They haven't released any


Obviously a balance would be best, but as someone who went to a very grade-inflated school, I do believe that grade inflation gets in the way of education substantially. When you can get through classes with very little effort and understanding and know you will get a sufficient grade, many people will simply not learn the material deeply.


The material outcome is what should be the goal. Tests are a relatively brute way to try and determine how well the student understands the material, but conversations about grade inflation and "back in my day getting a grade was hard", and professors purposefully putting difficult questions (not in content but in presentation of the question) all betray the inherent goal being pursued.

Its all Goodhart's law problem, but we are missing the forest for the trees talking about grades and tests when what we want is people to be educated, and critical thinkers and competent in their area and due to a comprehensive way to evaluate that we end up talking about grade inflation or how Yale vs Berkeley gives letters at the end of a semester


I mean, so many graded assignments are online now that very little technology is needed to cheat. I would guess that is the largest driver in increased cheating at universities.


As someone who attended an elite school in the post-covid era, here was my experience:

There is relatively little stigma against cheating. Maybe in smaller seminars and classes with higher collaboration there is some, but much less so in large STEM lectures. Many of the incentives in classes where exams were online led to arms races and widespread cheating (without exaggeration, over 80% of the class). For instance, a certain math class I knew of had all grades based on remote and often asynchronous tests. Many people would cheat/collaborate and ace them, leading to the professor increasing difficulty (as scores were very high). This led to more cheating and so on. It got to the point where the problem sets had such difficult problems in this intro class that only a handful of people (who had taken advanced course work in high school) in the entire 100+ person seminar were distributing proofs for everyone else. Really not great dynamics all around and it's worth noting that my school does not have a reputation for being ones with an especially competitive and cutthroat culture.


Sort of. Putting current knowledge into a number can be pretty interesting / useful though. Like many people, I read headlines and pay attention to what's happening in international politics, but from those it's hard to have any sense of how much reality there is to bluster in Iran/Panama/Venezuela/Greenland just from general discourse and media. For me, prediction markets have been very helpful in offering some sort of grounding beyond the general noise in areas where I have very little intuition or realistic sense of the possibilities.


I’m not so sure. From talking to some of my own friends at google they feel that antigravity/gemini models are handicapping them and would much rather be using claude code (which only deepmind gets to use)


Sure, but there's cavernous distance between "google = john deere" and "darn I have to use Gemini"


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: